Code and data for "When Does Skew-Symmetric Wedge Memory Help? Mapping the Boundary Between Global Aggregation and Associative Recall."
Tensor Cache (arXiv:2605.22884) found that writing a skew-symmetric (wedge-product) memory instead of a plain outer-product one hurts associative recall, and traced this to a crosstalk term in the read. That ablation was run in a causal, streaming setting. This repo tests whether the same failure shows up in a non-causal, global-aggregation setting, or whether it's specific to recall.
Short version of the result: skew-symmetric memory loses to the symmetric baseline on every recall configuration tested (60/60), but wins or ties on most aggregation configurations (22/30). The paper has the full writeup.
experiment.py— trains all four models (MeanMLP, Attention, Symmetric Memory, Exterior Memory) on both tasks, across two capacity settings, five seeds, and a size sweep. Writesresults_full.csv. Supports a--quickflag for a fast smoke-test run (see below).results_full.csv— raw output from the sweep, one row per run. This is what every table and figure in the paper is computed from.make_figures.py— regenerates the three paper figures from the CSV.Exterior-Memory.tex/Exterior-Memory.pdf— the paper.
pip install -r requirements.txt
python experiment.py
python make_figures.py
The sweep is 384 short training runs (15k steps each, small models). It splits across however many CUDA devices are visible; on a single GPU it just runs sequentially. On two T4s it took a few hours.
For a fast sanity check before committing to the full sweep, run
python experiment.py --quick, which cuts training to a couple hundred
steps on a single seed over a reduced size sweep. Run
python experiment.py --help to see all options (steps, seeds, output path).
results_full.csv already has everything. make_figures.py reads it
directly, so you can check the figures/tables against the paper without
running the sweep again.
@article{shabani2026exterior,
title={When Does Skew-Symmetric Wedge Memory Help? Mapping the Boundary
Between Global Aggregation and Associative Recall},
author={Shabani, Laerti},
journal={arXiv preprint},
year={2026}
}
MIT, see LICENSE.