Language Model 'Shape': designing architectures around agent workflows
Alex Zhang (independent)
Alex Zhang argues the input/output structure of language models has been frozen on the decoder-only Transformer since ChatGPT, with all agent innovation going into harnesses built around that fixed shape. He proposes designing model shapes for the task instead: agent trajectories need dense attention only on recent observations, so a hybrid recurrent-plus-Transformer shape could give an auto-compacting history without manual context folding. He points to Jev, a prefill-only model constrained to [0,1] outputs, as evidence that alternate tradeoff spaces exist, and argues open-weight distillation makes shape research newly feasible for independent researchers.
Why it matters
A concrete architectural program for agent-native models that challenges the one-size decoder-only default.
Importance: 2/5
Influential independent position paper on agent-native architectures