Blog 2026-10-07

The agent stopped. Did it answer the question?

An agent loop stopping is an execution event. It does not tell you whether the agent answered the question. A tool-call limit is useful for bounding work, but reaching it should not turn an incomplete investigation into a confident answer.

Hoyt Emerson’s “I Built My First AI Agent Loop on DuckDB (Python + Haiku)”, published by The Full Data Stack on October 5, prompted this example. Its description covers a database-query tool, iteration limits and trace inspection. The tutorial is sponsored by MotherDuck. We reviewed its description and chapters, not a loaded transcript, and did not reproduce the creator’s outcomes. His Python harness and trace database are separate from Bufflehead; no endorsement is implied.

One question, three evidence checks

Use fictional order 101 and ask: “How much captured cash remains after refunds?” The order total is 5,000 cents. Captured payments total 5,000 cents. Refunds total 1,200 cents. The known answer is therefore 3,800 cents, or $38.00.

Our deliberately sequential harness checks the order, captured payments and refunds in three read-only queries:

SELECT total_cents FROM orders WHERE id=101;
SELECT SUM(captured_cents) FROM payments WHERE order_id=101;
SELECT SUM(refunded_cents) FROM refunds WHERE order_id=101;

With a two-call budget, the harness has read the first two results but has not inspected refunds. It must not answer $50.00. That is captured cash before refunds, not the requested net amount.

budget: 2
calls: 2
status: budget_exhausted
remaining: [refunds]
net_cents: null

A useful incomplete response would say: “I verified $50.00 in captured payments, but have not checked refunds. The net amount remains unknown.” That response is illustrative wording, not output from a model.

With a three-call budget, all three checks run. The harness can compute the known answer:

budget: 3
calls: 3
status: complete
remaining: []
net_cents: 3800

What we tested

We ran this deterministic Python harness against a disposable in-memory SQLite fixture with query-only mode enabled. Each trace records the query purpose, SQL and returned rows. Both budget cases matched their expected results. Five additional checks passed: zero calls suppress the answer; one call leaves payment and refund checks outstanding; a four-call budget still stops after three completed checks; missing payment records suppress the answer even with three calls; and a write attempt is refused.

The verification script runs with Python’s standard library and creates only fictional, temporary data. This is a test of completion accounting, not a live AI comparison, a Bufflehead integration test or an accuracy benchmark.

More calls are not automatically better

The sequential plan is chosen to make the unfinished step visible. An agent could combine the two totals into one query and answer within a smaller budget. It could also spend many calls asking irrelevant questions. The number of calls is not a measure of correctness.

For a real model comparison, keep the question, fixture, model and supplied business definitions fixed. Run fresh sessions with two bounded budgets. Save their actual queries, tool results, stop reasons and answers, then compare each answer with the known $38.00 result. Do not claim a larger budget improved accuracy unless the observed runs support it. We have not run that model experiment here.

Completion belongs to the client

Record at least the stop reason, evidence obtained, checks still outstanding and answer status. Budget exhaustion should leave unknown fields unknown. Missing records, query errors and tool failures need their own status rather than being silently treated as zero.

Our fixture contains an explicit refund record. An empty refund result could mean “no refunds” or missing data in a real system; deciding which requires your data contract. Likewise, this example measures captured cash less refunds, not accounting revenue, profit or settlement balance.

Bufflehead supplies read-only database access through MCP. Your AI client or external harness controls tool-call budgets, records traces and decides whether enough evidence exists to answer. This example’s SQLite harness is external to Bufflehead; it does not claim that Bufflehead provides those controls.

Try Bufflehead with a known-answer dataset and inspect what the AI actually queried. When the loop stops, check what it learned before trusting what it says.

Try it on your own data.

Native binary, no Docker, no cloud egress. Free for local files and direct database connections — request a demo for the AWS SSM path.

Download free Request a demo