Beyond Vector Search: Building a Coding Agent That Actually Remembers (and Forgets)

by | Jul 1, 2026 | Articles | 0 comments

You tell your AI coding agent to build a CRUD API using FastAPI. It generates a beautiful, type-hinted masterpiece.

Two weeks later, you ask it to add an OAuth2 login flow.

Suddenly, it starts spewing Django middleware, settings.py configurations, and a manage.py file.

Why? Because your agent doesn’t actually remember your architectural constraints. It just retrieves the most statistically probable token based on your prompt—and statistically, “Django” is more popular than “FastAPI” on the internet.

Current AI coding agents suffer from a massive identity crisis. They rely on vector RAG (Retrieval-Augmented Generation) over codebases, which is essentially a giant, static filing cabinet. It has no concept of time, no understanding of contradiction, and—worst of all—it never forgets.

Today, we are open-sourcing the architecture for a Living Memory Layer that fixes this. It’s a two-pillar system: Contradiction-Aware Memory and Principled Forgetting.

Here is how we are building the agent that finally stops hallucinating imports and actually learns from its history.


The “Frankenstein Code” Problem

Before we dive into the solution, let’s diagnose the specific pain points of standard AI memory in a dev environment:

  • Context Poisoning: The agent remembers you used pytest three months ago, but you’re now writing unittest boilerplate. It mixes the two, generating an unholy test suite that works in neither.
  • The “Undead” File: You delete a legacy module, but the vector embedding persists. The agent keeps suggesting fixes for lines that haven’t existed for weeks.
  • Bug Recurrence: You fixed a tricky concurrency bug yesterday. Today, the agent suggests the exact same broken threading pattern because it forgot the context of the fix.
  • Dependency Amnesia: You upgraded from Python 3.9 to 3.12. The agent still suggests pkg_resources instead of the standard importlib.metadata.

Vector databases are passive. They don’t care about these conflicts. Our system is active.


Pillar 1: The Referee (Contradiction-Aware Memory)

We treat memory ingestion like a code review. Every time the agent writes a new diff, the user gives a new instruction, or a dependency updates, the system doesn’t just save it—it validates it against everything else stored.

We have built a Conflict Detection Engine (CDE) that checks for three specific types of contradictions:

  • Semantic Contradiction: We compare new user intents against old ones using embedding similarity. If you previously said, “Strictly use SQLAlchemy 1.4” and now say, “Use raw async PG driver,” the system flags a severe warning.
  • Structural Contradiction: We scan the Abstract Syntax Tree (AST). If the agent tries to write a function signature that collides with a stored helper function in a different file, or if it creates a circular import, it gets flagged.
  • Dependency Contradiction: The agent parses pyproject.toml and requirements.txt against its historical memory. If a new library requires a version that breaks an older stored logic, the system raises a red flag.

How it Resolves (The Workflow)

Not all conflicts require your input. We built a triage system:

  • Auto-Resolution: If the new memory has a timestamp newer than the old one, the agent assumes deprecation. It overwrites the old memory and logs the reasoning in a ResolutionLog.
  • Ask-User: For medium-severity conflicts (e.g., mixing testing frameworks), the agent pauses, highlights the two contradictory pieces of code in your IDE, and asks a single, clear multiple-choice question.
  • Block & Escalate: If a conflict involves a core security patch (e.g., Django 3.4 vs 5.0), the agent halts the task entirely. It refuses to write code until you explicitly choose a path.

Pillar 2: The Cleaner (Principled Forgetting)

Here is the counter-intuitive truth: To be intelligent, an AI must be allowed to forget.

Human memory is defined by decay. We forget the color of the car we drove yesterday because it’s irrelevant to driving today. Our agents need the same efficiency. If we let memory grow infinitely, the agent spends 90% of its context window on noise, drowning out the signal.

We built a Principled Forgetting Scheduler (PFS) that runs periodically (or during idle time). It calculates a Future Utility Score for every single memory node:

Score = (Recency * 0.3) + (Edit Frequency * 0.3) + (Centrality in Call Graph * 0.4)

Based on this score, the agent takes action:

  • Soft Delete (Archival): Memories with a score below 0.2 are compressed into a tiny “Cold Storage” summary. This reduces token overhead without permanently losing the thread.
  • Hard Delete: Memories related to deleted files or abandoned Git branches are purged permanently. If the file doesn’t exist, neither does the memory of it.
  • Deprecation Tag: If a memory refers to a library that is now 2 major versions behind, the agent automatically tags it as DEPRECATED and appends a note: “Consider using X instead—this logic is from 2024.”

Pillar 3: The Historian (The Episodic “Why” Graph)

This is the secret bonus feature.

We store not just what code was written, but why it was written that way.

When the agent retrieves a specific database connection logic, it also retrieves the decision thread attached to it: “Chose AsyncPG to handle 10k concurrent connections in 2025.”

If the current task involves a basic CRUD app with only 50 users, the agent sees the “Why” and might deliberately choose to simplify the pattern, referencing the old code only as a footnote. It stops copy-pasting enterprise patterns into weekend projects.


A Real-World Scenario

Let’s walk through a day in the life of this system:

  1. Morning: You tell the agent, “Migrate our requirements from pip to Poetry.”
  2. The Action: The agent runs the PFS. It archives all memories tied to requirements.txt and pip install commands. It promotes the new poetry.lock to “High Priority” memory.
  3. Afternoon: You ask the agent to fix a failing CI/CD pipeline.
  4. The Magic: The agent starts to write a fix. Its LLM wants to output pip install -r requirements.txt.
  5. The Intercept: Our pre-generation hook intercepts this. The Conflict Engine checks the active memory, sees the archived requirements.txt, and replaces it with poetry install in the output buffer.
  6. The Result: The agent writes perfect code without you ever having to correct it. It remembered to forget the old way.

Success Metrics (The Hard Numbers)

We are tracking this system against rigorous KPIs in our internal beta:

  • 40% Reduction in hallucinated imports and missing module errors in unit tests.
  • 50% Decrease in user correction interventions (users are saying “No, use the other package” half as often).
  • Context Optimization: We are maintaining active token usage at ≤ 5k tokens, even for projects that have been active for 6+ months.

The Roadmap (How We’re Building It)

We aren’t shipping this all at once. Here is our phased approach:

  • Phase 1: The Referee. Building the semantic and AST conflict detectors. Let the agent ask questions.
  • Phase 2: The Cleaner. Implementing the Utility Score and the archival cron job.
  • Phase 3: The Historian. Linking the “Why” graph to Git commit history and dependency scanners.
  • Phase 4: Multi-Agent Sync. Enabling this shared memory across a swarm of agents working on a monorepo, with Git-revert awareness.

Conclusion: The Disciplined Architect

The biggest bottleneck in AI software engineering isn’t the LLM’s reasoning ability—it’s the memory scaffolding surrounding it.

We are tired of agents that act like junior interns with amnesia. The future belongs to agents that act like disciplined senior architects: they challenge your contradictions, they remember the historical context, and crucially, they know when to let go of the past.

If you are building an AI coding tool, stop building a “bigger” memory. Build a smarter one.

Have thoughts on AI memory architectures? Drop us a comment below or check out our GitHub discussions.

Written by

Related Posts

0 Comments

Submit a Comment

Your email address will not be published. Required fields are marked *