I accidentally turned LLM memory into program analysis
An engineer using LLM agents for vulnerability research found that while the models are increasingly good at navigating large codebases and exploring attack surfaces, they lose track of established facts over multi-hour investigations: they suggest already-ruled-out approaches, forget false assumptions, and reason from stale observations. Telling an LLM something is wrong does not make it stop believing dependent conclusions.
Existing memory systems store past conversations or observations, embed them, and retrieve relevant pieces on demand, but the author wanted the model to maintain what is currently known, not just recall what was said. In the example, if attacker controls object_a, object_a points to object_b, and object_b is a kernel object, the model should conclude the attacker controls a kernel object; but if later LLDB shows object_a no longer points to object_b, retrieval may pull contradictory memories and the model must recompute which conclusions survive.
The author recognized this as a program analysis problem: facts and inference rules derive a fixed point of consequences, and when an input fact changes, incremental update techniques can adjust exactly the affected results instead of relying on the LLM to re-derive everything from a noisy memory subset.