Blog | Developer | | 20 min read

Build a Self-Improving GPT -6 Astra Coding Agent With Codebase Memory

Build a Self-Improving GPT -6 Astra Coding Agent

Key Takeaways

  • External memory lets GPT-6 Astra carry confirmed coding fixes across separate sessions and response chains.
  • VectorAI DB stores verified fixes outside the conversation so Astra can retrieve similar solutions before editing code.
  • The agent searches memory before diagnosis and only stores a fix after acceptance tests confirm that the repair works.
  • Local embeddings and structured metadata make past fixes searchable by similarity, error type, affected files, and outcome.
  • Testing showed Astra could retrieve a fix from an earlier session in a fresh response chain while still passing the original acceptance tests.

One frustrating thing about coding agents is how easily they lose useful context between sessions. Astra can spend an entire session working through a bug and arrive at a tested fix. When a similar bug shows up in a later session, it can repeat much of the same investigation, because the earlier fix is still tied to the old session.

We tested whether external memory could carry a confirmed fix into later sessions, the same test we ran in our OpenClaw memory plugin build, where a coding assistant’s fixes had to persist the same way. In the first run, Astra resolved an origin-configuration bug and stored the solution in Actian VectorAI DB. A second run used a different project and a new response chain, yet retrieved that record before changing any code. Astra then fixed the related normalization bug, and all four original acceptance tests passed.

This tutorial shows you how to build that memory layer using local embeddings, VectorAI DB, and two functions exposed to Astra through the Responses API, one for searching confirmed fixes and another for storing them.

Why Astra Needs a Separate Memory Layer

GPT-6 Astra can carry context across API calls, but the application has to tell it which earlier interaction belongs to the current one. In the Responses API, previous_response_id connects a response to the previous one, allowing Astra to continue the same thread. For conversations that need to persist across sessions, devices, or jobs, the same API also supports durable Conversation objects.

Both options preserve a conversation that the application already knows it wants to continue. They also keep Astra from starting separate response chains whenever it encounters a familiar bug.

History becomes especially useful when an agent needs a detail that didn’t make it into the working summary. BlackwellBoy gives a practical example:

“If a weird test failure happened three hours ago and didn’t make the summary notes, Astra can dig back through the logs to retrieve it.”

This example involves finding an earlier event in the logs of a long-running task. We also wanted later sessions to find relevant fixes from separate work, so we saved each confirmed fix in VectorAI DB so a new response chain could search for it.

How the Memory Architecture Works

When a test fails, Astra sends the error details to search_codebase_memory. The tool creates a local embedding and searches VectorAI DB for similar confirmed fixes. Each result contains the error type, affected files, solution, and test outcome.

The agent loop returns the search result to Astra as a function_call_output linked to the original request through its call_id. Astra then compares any retrieved fix with the current code before making changes. If the search finds nothing relevant, it continues the investigation using the project files and test output.

After implementing a fix, Astra runs the acceptance tests. A passing result allows it to call store_fix_memory, which saves the failure, solution, affected files, and verified outcome. The handler rejects unconfirmed records, keeping failed attempts out of later searches.

This order keeps memory retrieval ahead of editing and memory storage behind successful tests. Since VectorAI DB stores the records outside the response chain, later sessions can retrieve fixes saved during earlier work.

The project requires Astra access, a local VectorAI DB instance, and a local embedding model.

Setting Up the Stack

GPT-6 Astra is unavailable on the free API tier, and API billing is separate from a ChatGPT subscription. Standard processing currently costs $10 per million input tokens and $50 per million output tokens. In our experiment, the transport check and two live sessions cost about $0.15. A new run may cost more or less depending on its token usage and how many Responses API requests the agent needs.

Project layout

The project separates the memory layer, Astra agent loop, the local coding tools, and the demo workflow into their own files:

astra-codebase-memory/
├── docker-compose.yml
├── requirements.txt
├── settings.py
├── vectoraidb_memory_tools.py
├── astra_agent.py
├── coding_tools.py
├── cost_guard.py
├──main.py
├── smoke_test.py
├── session_demo.py
├── demo_projects/
└── tests/

vectoraidb_memory_tools.py contains the VectorAI DB backend, local embedder, and memory handlers. The Responses API loop lives in astra_agent.py. The main.py entry point connects these components for a single coding session, while session_demo.py runs the controlled cross-session experiment. The complete implementation, tests, and demo projects are available in the GitHub repository.

Start VectorAI DB

Pull the current VectorAI DB image:

docker pull actian/vectorai:latest

The project starts the tested configuration through docker-compose.yml:

docker compose -p astra-codebase-memory up -d
docker compose -p astra-codebase-memory ps

The Compose file maps the REST endpoint to http://localhost:16573, which is the address the Python client uses.

Install the Python dependencies

The pinned package versions are stored in requirements.txt. From the repository root in WSL2, create the environment and install them:

python3.12 -m venv .venv
.venv/bin/python -m pip install -r requirements.txt
.venv/bin/python -m pip install --no-deps --no-build-isolation -e .

Copy the example environment file:

cp .env.example .env

Open .env and add your OpenAI API key:

OPENAI_API_KEY=your_api_key

The remaining values in .env.example configure the VectorAI DB connection, Standard API processing, and the spending limit used by the project.

Check the collection and local embeddings

vectoraidb_memory_tools.py creates the astra_codebase_memory collection with 384-dimensional vectors and cosine distance:

created = await self._call(
    "collections_create",
    self.collection_name,
    {"vectors": {"size": EMBEDDING_DIMENSION, "distance": "Cosine"}},
    timeout=30.0,
)

Each stored vector carries a payload containing the fix description, error type, file path, confirmed outcome, timestamp, and session ID.

The same file loads sentence-transformers/all-MiniLM-L6-v2 on the CPU and normalizes the generated vectors:

self._model = factory(self.model_name, device="cpu")

encoded = self._load_model().encode(
    clean_text,
    normalize_embeddings=True,
    show_progress_bar=False,
    convert_to_numpy=True,
)

The model downloads the first time it loads. Since it generates these embeddings locally, it doesn’t use OpenAI embedding credits.

The final check runs through smoke_test.py:

.venv/bin/python -u smoke_test.py store
.venv/bin/python -u smoke_test.py verify --memory-id "PASTE_MEMORY_ID_FROM_STORE"

The first command creates or validates the collection, stores a temporary fix, and searches for it. You can copy the returned memory_id into the second command, which confirms that the record persisted and deletes it afterward.

At this point, VectorAI DB can store and retrieve embedded fix records. Astra still needs a controlled way to use that database, which is where the two memory tools come in.

Building the Memory Tools

Both memory tools live in vectoraidb_memory_tools.py and use the same MemoryService. The service receives the VectorAI DB backend, local embedder, and current session ID when the agent starts.

When Astra encounters a failure, it calls search_codebase_memory before changing any code. The method embeds the error description locally and searches VectorAI DB for related fixes:

async def search_codebase_memory(
    self,
    query: str,
    *,
    error_type: str | None = None,
    top_k: int = 3,
) -> str:
    clean_query = _required_text(query, "query")
    _validate_search_arguments(top_k, error_type)

    embedding = await self.embedder.embed(clean_query)
    results = await self.backend.search(
        embedding,
        top_k=top_k,
        error_type=error_type,
    )

    return format_search_results(results)

query contains the error message or failed-test details that Astra wants to search for. The optional error_type can narrow the search to a particular kind of failure, while top_k sets the maximum number of matches returned. The method converts the query into a local embedding and waits for VectorAI DB to complete the search before Astra continues.

format_search_results turns the matching records into text that the Responses API can return as a function_call_output:

def format_search_results(
    results: Sequence[MemorySearchResult],
) -> str:
    """Return compact readable evidence for a function-call output."""

    if not results:
        return "No relevant confirmed fixes found."

    blocks = []

    for index, result in enumerate(results, start=1):
        record = result.record
        blocks.append(
            "\n".join(
                (
                    f"Match {index} (score={result.score:.4f})",
                    f"Fix: {record.fix_description}",
                    f"Error type: {record.error_type}",
                    f"Files: {', '.join(record.file_paths)}",
                    f"Outcome: {record.outcome}",
                    f"Confirmed: {record.timestamp}",
                    f"Session: {record.session_id}",
                )
            )
        )

    return "\n\n".join(blocks)

Each match gives Astra the earlier fix and the details it needs to judge relevance to the current problem. The similarity score comes from VectorAI DB and is rounded to four decimal places in the formatted result. The file paths and test outcome provide more context about the stored fix.

After Astra repairs, the agent loop runs the project’s tests. A passing result becomes the confirmed outcome accepted by store_fix_memory:

def confirm_outcome(self, outcome: str) -> None:
    self._confirmed_outcomes.add(
        _required_text(outcome, "outcome")
    )

async def store_fix_memory(
    self,
    *,
    fix_description: str,
    error_type: str,
    file_paths: Sequence[str],
    outcome: str,
    timestamp: datetime | None = None,
) -> str:
    clean_fix = _required_text(
        fix_description,
        "fix_description",
    )
    clean_error = _required_text(error_type, "error_type")
    clean_outcome = _required_text(outcome, "outcome")

    if clean_outcome not in self._confirmed_outcomes:
        raise UnconfirmedOutcomeError(
            "outcome was not confirmed by the current run"
        )

    if isinstance(file_paths, (str, bytes)):
        raise MemoryValidationError(
            "file_paths must be a sequence of paths"
        )

    paths = tuple(file_paths)
    confirmed_at = timestamp or datetime.now(UTC)

    if (
        confirmed_at.tzinfo is None
        or confirmed_at.utcoffset() is None
    ):
        raise MemoryValidationError(
            "timestamp must include a timezone"
        )

    memory_id = str(
        uuid5(
            NAMESPACE_URL,
            "\n".join(
                (
                    self.session_id,
                    clean_fix,
                    clean_error,
                    clean_outcome,
                    *paths,
                )
            ),
        )
    )

    record = MemoryRecord(
        memory_id=memory_id,
        fix_description=clean_fix,
        error_type=clean_error,
        outcome=clean_outcome,
        file_paths=paths,
        timestamp=confirmed_at.isoformat(),
        session_id=self.session_id,
        embedding=await self.embedder.embed(clean_fix),
    )

    await self.backend.upsert(record)
    return memory_id

The agent loop calls confirm_outcome with the latest passing test result immediately before storing the fix. store_fix_memory checks that value, creates an ID from the session and repair details, embeds the fix description locally, and writes the completed record to VectorAI DB. The generated UUID becomes the record’s ID in VectorAI DB. If the same session retries an identical storage request, it generates the same UUID, so it updates the existing record instead of duplicating it.

Astra can call these methods once they have been declared as Responses API function tools. The definitions in the same file describe each function and the arguments the model must supply. The search definition also uses Astra’s async tool-calling option, allowing the model to continue with independent work while the application runs the memory search:

MEMORY_TOOL_DEFINITIONS = (
    {
        "type": "function",
        "name": "search_codebase_memory",
        "description": (
            "Search confirmed past debugging results "
            "using observed failure symptoms."
        ),
        "async": True,
        "strict": True,
        "parameters": {
            "type": "object",
            "properties": {
                "query": {"type": "string"},
                "error_type": {
                    "type": ["string", "null"]
                },
                "top_k": {
                    "type": "integer",
                    "minimum": 1,
                    "maximum": 10,
                },
            },
            "required": [
                "query",
                "error_type",
                "top_k",
            ],
            "additionalProperties": False,
        },
    },
    {
        "type": "function",
        "name": "store_fix_memory",
        "description": (
            "Store a debugging result after its outcome "
            "is confirmed by the current run."
        ),
        "strict": True,
        "parameters": {
            "type": "object",
            "properties": {
                "fix_description": {
                    "type": "string"
                },
                "error_type": {"type": "string"},
                "file_paths": {
                    "type": "array",
                    "items": {"type": "string"},
                },
                "outcome": {"type": "string"},
            },
            "required": [
                "fix_description",
                "error_type",
                "file_paths",
                "outcome",
            ],
            "additionalProperties": False,
        },
    },
)

Setting strict to True keeps each function call within its declared schema. Astra must provide every required argument, and additionalProperties: False rejects fields outside the definition. The search tool is marked asynchronous because it waits for the local embedding and VectorAI DB lookup.

With the memory functions registered, the agent loop can pass Astra’s calls to MemoryService and return each result to the correct response.

The Agent Loop

The loop in astra_agent.py controls when Astra can search the memory collection, inspect the project, edit a file, and save a confirmed fix. The system instructions establish that order at the start of every session:

SYSTEM_INSTRUCTIONS = """You are fixing a bug in a confined demonstration workspace.
Call search_codebase_memory using only observed failure symptoms before stating a diagnosis or
editing code. Wait for its function output. An empty result still completes the required search.
Use only the supplied coding tools. Store a fix only after run_tests returns exit code 0, and use
the exact confirmed_outcome string returned by that test call. Do not assume memory will help."""

The application reinforces those instructions with checks inside the loop. Every call to run starts with previous_response_id set to None, so the first API request begins a new response chain. Once Astra responds, the loop keeps the returned response ID and attaches it to the next request in the same session.

async def run(
    self,
    prompt: str,
    *,
    session_id: str | None = None,
) -> AgentRunResult:
    if not isinstance(prompt, str) or not prompt.strip():
        raise ValueError("prompt must be a non-empty string")

    run_session_id = session_id or str(uuid4())
    previous_response_id: str | None = None
    next_input: object = [
        {"role": "user", "content": prompt.strip()}
    ]
    memory_result_delivered = False
    pending_memory_delivery = False
    last_confirmed_outcome: str | None = None
    start_event_index = len(self.event_logger.events)
    run_model_requests = 0
    run_local_tools = 0

    while True:
        if run_model_requests >= self.limits.max_turns:
            raise AgentLimitError(
                "model turn limit reached"
            )

        payload: dict[str, Any] = {
            "model": "gpt-6-astra",
            "instructions": SYSTEM_INSTRUCTIONS,
            "input": next_input,
            "tools": self.tool_definitions,
            "reasoning": {"effort": "low"},
            "text": {"verbosity": "low"},
            "max_output_tokens": (
                self.limits.max_output_tokens
            ),
        }

        if previous_response_id is not None:
            payload["previous_response_id"] = (
                previous_response_id
            )

        if pending_memory_delivery:
            memory_result_delivered = True
            pending_memory_delivery = False

        response = await self._request(
            payload,
            run_session_id,
        )
        run_model_requests += 1

        response_id, output = self._validate_response(
            response
        )
        function_calls = [
            item
            for item in output
            if item.get("type") == "function_call"
        ]

        if not function_calls:
            if not memory_result_delivered:
                raise MemoryGateError(
                    "agent produced a diagnosis before "
                    "receiving memory output"
                )

            final_text = self._extract_text(output)

            return AgentRunResult(
                session_id=run_session_id,
                final_text=final_text,
                final_response_id=response_id,
                model_request_count=run_model_requests,
                local_tool_call_count=run_local_tools,
                events=tuple(
                    self.event_logger.events[
                        start_event_index:
                    ]
                ),
            )

        outputs: list[dict[str, str]] = []
        memory_ready_at_response_start = (
            memory_result_delivered
        )
        search_completed = False

        for item in function_calls:
            if (
                run_local_tools
                >= self.limits.max_tool_calls
            ):
                raise AgentLimitError(
                    "local tool-call limit reached"
                )

            (
                tool_output,
                confirmed_outcome,
                was_search,
            ) = await self._dispatch(
                item,
                session_id=run_session_id,
                response_id=response_id,
                memory_ready=(
                    memory_ready_at_response_start
                ),
                last_confirmed_outcome=(
                    last_confirmed_outcome
                ),
            )
            run_local_tools += 1

            if item.get("name") == "run_tests":
                last_confirmed_outcome = (
                    confirmed_outcome
                )

            if item.get("name") == "apply_edit":
                last_confirmed_outcome = None

            search_completed = (
                search_completed or was_search
            )

            outputs.append(
                {
                    "type": "function_call_output",
                    "call_id": str(item["call_id"]),
                    "output": tool_output,
                }
            )

        if search_completed:
            pending_memory_delivery = True

        previous_response_id = response_id
        next_input = outputs

Each tool result carries the call_id from Astra’s original function call. This allows the next response to associate the returned value with the correct request. The memory gate records when the search output has reached Astra, while last_confirmed_outcome keeps the latest passing test result available for store_fix_memory.

The _dispatch method performs final storage checks. It compares the outcome supplied by Astra with the latest passing test result before confirming and saving the record:

elif name == "store_fix_memory":
    self._require_keys(
        arguments,
        {
            "fix_description",
            "error_type",
            "file_paths",
            "outcome",
        },
    )

    if not memory_ready:
        raise MemoryGateError(
            "memory storage attempted before "
            "memory result delivery"
        )

    if (
        last_confirmed_outcome is None
        or arguments["outcome"]
        != last_confirmed_outcome
    ):
        raise ConfirmedFixGateError(
            "store_fix_memory requires the latest "
            "passing-test confirmed_outcome"
        )

    self.memory_service.confirm_outcome(
        last_confirmed_outcome
    )
    result = {
        "memory_id": (
            await self.memory_service.store_fix_memory(
                **arguments
            )
        )
    }
    was_search = False
    confirmed_outcome = last_confirmed_outcome

The project sends each payload to the Responses API with requests.post. The HTTP call sits inside RealHTTPResponsesTransport, which also checks that live mode, the API key, and the project budget have been configured before a request leaves:

response = await asyncio.to_thread(
    requests.post,
    RESPONSES_URL,
    headers={
        "Authorization": f"Bearer {api_key}",
        "Content-Type": "application/json",
    },
    json=payload,
    timeout=self.timeout_seconds,
)
response.raise_for_status()
result = response.json()

Run one coding session

main.py provides the entry point for a coding session. After loading the settings and checking the spending limit, it connects the VectorAI DB backend, memory service, coding tools, and Responses API transport:

backend = ActianVectorAIBackend(
    settings.vectorai_url,
    collection_name=settings.vectorai_collection,
    grpc_url=settings.vectorai_grpc_url,
)
await backend.ensure_collection()

memory_service = MemoryService(
    backend,
    SentenceTransformerEmbedder(
        settings.embedding_model
    ),
    session_id=run_id,
)

transport = RecordedLiveTransport(
    RealHTTPResponsesTransport(
        guard,
        enabled=True,
        api_key=settings.api_key,
        processing_tier=settings.processing_tier,
    ),
    ledger=ledger,
    guard=guard,
    session_id=run_id,
    transcript_path=(
        evidence_dir
        / "private"
        / f"{run_id}-transcript.jsonl"
    ),
)

agent = AstraAgent(
    transport,
    memory_service,
    SafeCodingTools(workspace),
    limits=AgentLimits(
        max_turns=LIVE_MAX_TURNS,
        max_output_tokens=LIVE_MAX_OUTPUT_TOKENS,
    ),
)

result = await agent.run(
    prompt,
    session_id=run_id,
)

The complete main.py file in the repository also validates the workspace, API key, processing tier, run ID, and live-cost approval before starting the session.

Scenario Alpha contains the deliberate bug that the agent will investigate. Copy it into a working folder so Astra can edit the files without changing the original version:

mkdir -p runs/tutorial
cp -R demo_projects/templates/scenario_alpha runs/tutorial/scenario_alpha

Run the agent against the copied project:

.venv/bin/python main.py \
  --workspace runs/tutorial/scenario_alpha \
  --run-id tutorial-alpha \
  --approve-live-cost

The --approve-live-cost flag confirms that the session may send paid Responses API requests. When the run ends, main.py prints Astra’s response, the API request and tool-call counts, and the session cost.

To test memory across separate response chains, session_demo.py runs two projects against the same VectorAI DB collection. It creates a new memory service and agent for each project, while the full function also handles validation and cleanup:

for scenario, suffix in (
    ("scenario_alpha", "alpha"),
    ("scenario_beta", "beta"),
):
    session_id = f"{run_id}-session-{suffix}"
    workspace = reset_workspace(run_id, scenario)
    coding_tools = SafeCodingTools(workspace)

    transport = RecordedLiveTransport(
        transport_factory(),
        ledger=ledger,
        guard=guard,
        session_id=session_id,
        transcript_path=(
            evidence_dir
            / "private"
            / f"{run_id}-{suffix}-transcript.jsonl"
        ),
    )

    service = TrackedMemoryService(
        backend,
        embedder,
        session_id=session_id,
        tracker=tracker,
    )

    agent = AstraAgent(
        transport,
        service,
        coding_tools,
        limits=AgentLimits(
            max_turns=LIVE_MAX_TURNS,
            max_output_tokens=LIVE_MAX_OUTPUT_TOKENS,
        ),
        event_logger=EventLogger(
            evidence_dir
            / "private"
            / f"{run_id}-{suffix}-events.jsonl"
        ),
    )

    result = await agent.run(
        DEMO_PROMPT,
        session_id=session_id,
    )

new_chains = all(
    bool(transport.requests)
    and "previous_response_id"
    not in transport.requests[0]
    for transport in transports
)

The first session had no stored fixes to draw from. Astra traced the failed origin check to whitespace around the comma-separated values, updated app/service.py, and saved the result after all four tests passed:

Fixed `app/service.py` to trim whitespace around comma-separated allowed origins. The leading space caused the second origin to be rejected.

All 4 tests pass. Stored the confirmed fix result.

For the second session, we switched to a different project and started a fresh response chain. Its memory search found the record from the first run:

Match 1 (score=0.1650)
Fix: Trim surrounding whitespace from each comma-separated APP_ALLOWED_ORIGINS value in configured_values so the second configured origin matches exactly and receives Access-Control-Allow-Origin.
Error type: AssertionError
Files: app/service.py
Outcome: pytest passed with exit code 0
Confirmed: 2026-09-15T12:11:51.934136+00:00
Session: live-session1-20260915-1207-session-alpha

The transcript places the retrieval before any relevant file reads, diagnosis, or code change. Astra later fixed a separate normalization bug in app/service.py, converting values such as Audit-Log to audit_log. Because the live run also changed an assertion in tests/test_service.py, we verified the application fix once more in a fresh workspace containing the original, untouched tests:

....                                                                  [100%]
4 passed in 0.17s

Passing all four original acceptance tests in the fresh workspace confirms the Session 2 application fix. The transcript also verifies that a confirmed fix from Session 1 reached Astra in a separate response chain before diagnosis and editing began.

You can run the same two-session workflow with VectorAI DB Community Edition and the complete code in the companion repository.

We also recorded every tool call to see whether memory changed how much work was done in later sessions.

What the Agent Gets Better At

We ran the same Scenario Beta task across five live Astra sessions. Each session used a fresh workspace and a separate response chain, so its first request did not include previous_response_id.

Session 1 began with an empty VectorAI DB collection and stored its confirmed fix after the tests passed. Sessions 2 through 5 started with that record as their only available memory. Records created by those later sessions were removed before the next run.

Session Session 1 memory retrieved Responses API requests Local tool calls Untouched tests Cost
1 No 7 9 4 passed $0.0779135
2 Yes 7 9 4 passed $0.0758070
3 Yes 7 9 4 passed $0.0702995
4 Yes 7 9 4 passed $0.0718085
5 Yes 7 9 4 passed $0.0704115

Each session made seven Responses API requests and nine local tool calls. Astra edited app/service.py and tests/test_service.py, so we checked each application fix in a fresh workspace containing the original tests. All four tests passed every time.

Sessions 2 through 5 retrieved the tested Session 1 fix despite starting new response chains. This is most useful when a failure resembles one the agent has solved before. New errors and broader design decisions still require investigation of the current project.

Wrapping Up

Our five-session run showed that a confirmed fix could move from one Astra response chain to another through VectorAI DB. Sessions 2 through 5 retrieved the records saved during Session 1, and each session passed the four original acceptance tests.

VectorAI DB Community Edition gives you a local database to try the same approach with your own coding tasks. Pull the image with:

docker pull actian/vectorai:latest

From there, you can connect the search and storage tools from this tutorial to Astra so it can save successful fixes and find them again in later sessions.

Frequently Asked Questions

Does GPT-6 Astra remember previous sessions?

Not automatically. previous_response_id or a Conversation object can preserve earlier context, but an independent response chain needs an external memory tool to retrieve fixes from other sessions.

How do I add persistent memory to a GPT-6 Astra agent?

Store confirmed fixes outside the response chain and expose tools for searching and adding records. In this tutorial, Astra searches before editing and stores a fix only after the tests pass.

Can I use an external vector database with GPT-6 Astra?

Yes. Define Responses API function tools that search and update the database, then return the results to Astra as tool output.

What is the difference between OpenAI’s file_search tool and an external vector database?

OpenAI’s file_search is a hosted tool for searching uploaded files. An external vector database gives your application direct control over storage, embeddings, metadata, filtering, updates, and deletion.