Back to the journal

The small-agent playbook.

A practical research plan for building agent experiments you can actually reproduce.

Research plan. These experiments have not been run; no measured results are claimed.

Start with one question

Does a smaller, more focused toolset help an agent complete a bounded repository task? That is a question we can investigate. “Which agent is best?” bundles too many decisions into a single answer. Start with a specific task, a fixed repository revision, and a result another person can inspect.

Our proposed first experiment compares the same agent with two tool configurations. The model, task instructions, source files, budget, and stopping rules stay fixed. This keeps the comparison focused, but model variability and service changes still need to be recorded. Alternate the order of conditions and repeat both before interpreting any difference.

Write the protocol before the prompt

  • Choose tasks with observable acceptance criteria and keep a separate set for evaluation.

  • Record model identifier, prompt version, repository commit, tool definitions, and environment.

  • Define what counts as success, failure, timeout, and manual intervention before running.

  • Keep failed attempts and full tool traces alongside successful ones.

A passing test suite is useful evidence, but it may miss unintended behavior. Add a review of the actual patch and report the limitations of each check.

Keep a run manifest

{
  "protocol": "small-agent-v1",
  "repositoryCommit": "<record exact commit>",
  "model": "<record model identifier>",
  "taskSet": "<versioned task file>",
  "toolConfiguration": "focused",
  "result": "not-run"
}

Repeat each condition, preserve raw observations, and publish the execution scripts. Record cost and elapsed time as observations rather than treating them as interchangeable with quality.

What we would publish

The deliverable should include the protocol, task set, manifests, traces with secrets removed, evaluation code, and a plain-language account of failures. We have not executed this study. A useful result may be that the narrower toolset makes no reliable difference.

For a primary reference on defining data and grading criteria, see the evaluation guide below. Our proposed protocol is editorial guidance; it is not a claim about a provider’s benchmark results.

OpenAI: working with evals

How this story was made

Prepared with AI assistance. Sources are linked in the story. This is a research plan, not a report of completed experiments. No measured results are claimed.

The conversation

MODERATED, ALWAYS

A useful question, a different result, a missing detail. Start there.