Afterimage

Research note · started October 2026

I set out to build a product and ended up with a question.

Who is writing this? About me →

I wanted an assistant that understands my screen in real time. I worked through how to build it, and concluded that combining existing methods and adding simple ideas would not get there. So I turned to studying and researching the part I could not see past. This page is that path, written down in order so it can be checked.

It started with the Claude for Startups program. I pulled a shelved project back out to apply with it, was declined automatically within minutes, and then took an honest look at whether the project deserved to exist. It did not. Section 01 has the details.

I do not know yet what the right first step is. My plan is to try anything connected to the question, small things first, and to treat this as an adventure.

Why "Afterimage", if the product is gone? It was the product's working name: what passes across the screen leaves a trace, a text journal, and the model answers from that trace instead of watching. I dropped the product and kept the name, because it fits the question better. An afterimage is what stays in the eye after the light is gone. It is not the thing itself, but what the eye made of it. Inside a language model, the afterimage is what remains after it reads an input: the representation it built, which no one can yet see directly. That is what I want to learn to look at.

When a language model reads the same content as prose, HTML, Markdown or a compact notation, where inside the model do those versions become the same thing, and does an unfamiliar format cost extra layers to decode?

  1. How this started
  2. The pain that stayed
  3. An earlier experiment, and how it corrected me
  4. Why can't a model just watch?
  5. Using the limits instead of fighting them
  6. Which format is cheapest to read?
  7. Adding it up: not enough
  8. Why can't we just look inside?
  9. Why I turned to research
  10. The question
  11. What I will do first
  12. What I bring, and what I don't

01 · How this started

Most days I have several VS Code windows open, several Claude Code sessions, and one Codex session, each on a different task. Moving context between them is copy and paste.

So I designed a desktop app to put them in one window: one task per panel, one git worktree per task, selective hand-off between agents, and an approval gate in front of every file write and shell command. I wrote a full design document and a working skeleton (about 2,200 lines of Rust, 50 tests).

Then I stopped. When I was honest with myself, I would not use it: the window juggling is annoying but tolerable.

A few weeks later I came across the Claude for Startups program and thought the shelved project might be a way in. I rewrote the pitch and applied. The answer came back within minutes: not approved, because the company could not be verified. That was fair. There is no company. I had written "solo, pre-incorporation" myself, and my website is a personal portfolio.

The rejection made me look harder at the project itself, this time as a product. Several tools, most of them free or open source, already run multiple coding agents in parallel, each in its own git worktree: Orca, Superset, Conductor, Vibe Kanban. The company behind Vibe Kanban, one of the closest matches, shut down in April 2026, and the project now lives on as a community effort. There was no realistic business in it, and I still would not use mine. A tool its own author would not use is not worth finishing, so I shelved it for good.

What survived the shelving was not the app but an itch underneath it.

02 · The pain that stayed

What I could not shrug off was something else. When an agent works on anything visual through a tool interface, the loop is look, fix, look, fix. Each look is a new request that pays for the whole context again. It is slow, it is expensive, and the result often falls short of what the underlying program can actually do, because the interface is a thin wrapper around an API.

The same thing shows up outside coding. I want an assistant that sits beside my screen and, when I ask "what is this?", already knows what I am looking at. Screen sharing gets close, but every moment of understanding is a separate, fully priced request. It is not real time in any sense I care about.

That became the goal: an assistant that keeps following what is on my screen, answers about it the moment I ask, and can lay a translation over text as it appears, at a cost low enough to leave running all day. The rest of this page is what happened when I tried to work out how to get there.

03 · An earlier experiment, and how it corrected me

Before all this I had run a small set of experiments on one idea: give the model the exact state a program already knows (coordinates, colors, a component tree) instead of a screenshot, and let it act through a compact notation instead of REST calls.

Two lessons stuck. I do not trust my own tests of my own ideas; comparisons go to an independent, blind run. And I look for prior work before I design, not after.

04 · Why can't a model just watch?

A transformer computes only when input arrives. Streaming interfaces keep a session open and feed it frames, so "a new request every time" is partly solved. What is not solved is cost: each frame costs tens to a few hundred tokens, services typically sample around one frame per second, and a long session fills even a very large context window. Deciding when something on screen is worth reacting to is still a research problem.

A person handles this with cheap peripheral vision and expensive focus. A model spends full price on every frame. I cannot change the model. I can only change what it is given.

05 · Using the limits instead of fighting them

That led to a design that leans on how transformers behave rather than against it:

That is the outline of a product. Before building it, I tried to work out how far it could go.

06 · Which format is cheapest to read?

If the journal is text, what kind of text? My first instinct was to invent the most compact language I could. The experiment above already argued against it. A model reads best what it saw most in training: prose, code, HTML, Markdown. A new notation needs a manual, and the manual is tokens.

There is a second reason. A model spends roughly the same computation on every token. Squeezing more information into fewer tokens lowers the bill and also lowers the computation available to think about that information. Some of the redundancy in natural language is waste; some of it is room to think.

So the cost of a format shows up in two places. One is the token count, which is on the bill. The other is how much of the model's depth goes into decoding the format before it can reason about the content. That second cost is never billed. It shows up as wrong answers. From outside, I can only see it indirectly, as accuracy per token.

07 · Adding it up: not enough

Put together, the levers I could find from outside the model are these. The figures are from published comparisons and price lists, not my own measurements yet.

Optimistically that is one to two orders of magnitude. That is real, and I may still build it. But it does not reach the goal. The model would still not follow the screen; it would only be asked about it more cheaply.

Every version of the thought experiment ended in the same place. Cheaper perception needed a cheaper representation. Choosing a representation needed knowing what the model actually does with it. And that happens inside the model, where I could not see. Stacking existing methods and simple ideas on top of each other did not get past that point. They only moved me closer to it.

08 · Why can't we just look inside?

My honest reaction was impatience. If the model stores everything as vectors, why not draw the map and read it? It turns out maps are being drawn. Anthropic decomposed a production model's activations into millions of interpretable features, and by amplifying one of them made the model believe it was the Golden Gate Bridge. Later work traced which features feed which on the way to an answer, and found the same concept used across languages in the middle of the model.

What I had underestimated is that a map is not an understanding:

The work is slow because the simple approaches were tried first and failed, not because nobody thought of them.

09 · Why I turned to research

So I changed course, from building to studying. I want to be clear about what that does and does not buy me.

It does not hand me the product. Even if I learn exactly how a model decodes a format, I cannot change Claude. Only the people who build it can, and results from small open models will not transfer for free.

What it buys is understanding instead of guessing. Every idea I had so far was a guess about what happens inside the model, checked only by its effect on the outside. I would rather learn to look than keep stacking tricks on a guess. If the research finds something, it shapes the input design first, and the product idea is parked, not dropped. And it is the only route I can see from where I am to the people who can change the model.

10 · The question

When a language model reads the same content as prose, HTML, Markdown or a compact notation, where inside the model do those versions become the same thing, and does an unfamiliar format cost extra layers to decode?

Why I think it is worth asking:

11 · What I will do first

  1. Learn the tools. Work through the open ARENA curriculum: build and train a small transformer, inspect it with TransformerLens, find features with sparse autoencoders, steer with them.
  2. Reproduce a known result before claiming a new one: find induction heads in a small model and watch them form during training.
  3. Run the format experiment on small open models: the same facts written in each format, the layer at which their representations converge, and whether a novel notation shifts that layer.
  4. Publish here, including results that go against me, with code and data.

I also want to check what I find in small open models against a frontier model's behavior from the outside, as accuracy per token across formats.

How I will work

I have no API budget, so the work runs on two tools that come with a paid Claude plan:

Today they are two separate apps. A research loop moves between them constantly: an experiment in one becomes a tool in the other, and the tool feeds the next experiment. I hope Anthropic will one day connect the two, so the same session can build and experiment with one shared context and one provenance record. I would use that every day.

12 · What I bring, and what I don't

I have built neural networks from scratch, forward pass, backward pass and optimizer, in a live browser lab at 7xz.dev/lab. I have run blind comparisons with independent agents and published the result that went against me. I write down what I do not know.

I have no lab affiliation, no publications and no compute budget. This page is the first step toward changing that.

Contact: me@7xz.dev

Sources

  1. Scaling Monosemanticity (Anthropic, 2024)
  2. Golden Gate Claude (Anthropic, 2024)
  3. On the Biology of a Large Language Model (Anthropic, 2025)
  4. In-context Learning and Induction Heads (Anthropic, 2022)
  5. Softmax Linear Units (Anthropic, 2022)
  6. Language models can explain neurons in language models (OpenAI, 2023)
  7. The structure of the nervous system of C. elegans (White et al., 1986)
  8. TOON, a compact serialization format for LLMs
  9. ARENA, an open curriculum for mechanistic interpretability
  10. Claude Science overview
  11. Vibe Kanban shutdown notice (bloop, April 2026)
  12. DirectShell: the accessibility layer as an app interface, with token comparisons
  13. Claude API pricing, including prompt caching
  14. Gemini API video understanding: frame sampling and tokens per frame