Skip to content

Context size

The selected context and provider guide covers source packets, exact raw-output recovery, explicit model profiles and optional candidate workers. These commands are part of the published CLI; inspect the installed version and each command’s help before use.

Before a task, estimate known text selected through files or explicit stdin with the bee token estimate command. The token preflight guide walks through comparing context selections, repeating the same input with --calls, allocating output and preserving totals with --limit 0. Requires Bee 1.0.0-beta.16 or later.

Select context, estimate it, review its adequacy, then perform the task separately. A smaller selection alone proves neither quality nor savings. Afterwards, use the measured-usage workflow below for an explicit supported log; preflight counts do not establish actual output, reasoning, cache use or provider consumption.

Terminal window
bee budget "catalog validation" --path src/Catalog/Product.cs --by-item --json

Preview the context Bee would render for a sample prompt. --by-item lists each rule, memory and other item by its actual UTF-16 character contribution. Section headings and wrappers are accounted for separately. Item content is not repeated. This command requires Bee initialization and a configured MongoDB.

Token figures use ceil(UTF-16 characters / 4). They are estimates, not model usage or billing records. Separately rounded item estimates may not add up to the total. Previewing a budget does not create a delivery receipt.

Terminal window
bee context audit --from AGENTS.md --against docs --root . --json

Compare an explicit Markdown source with an explicit file or directory. The audit reports hashes, sizes, source-section locations and normalized duplicate candidates. It runs offline before configuration or MongoDB access and changes no source files.

The source inventory lists up to 200 sections, largest body first, even when no duplicates exist. Body counts use normalized text; headings and boundary characters are accounted for separately from the original file size.

Literal blocks of at least 64 normalized characters are duplicate candidates. BOM and line endings are normalized; paragraph boundaries and ATX headings are excluded from matching. Other whitespace, code, case and punctuation are preserved. A match does not establish that an obligation can safely be removed.

Reads and output are bounded. Paths must stay inside the selected root; symbolic links and nonregular files are rejected. Unreadable, changed or over-limit input produces an incomplete report. On macOS use the physical path, such as /private/tmp, instead of the /tmp symbolic link.

Exit codes: 0 complete with no candidates; 1 complete with candidates; 4 invalid arguments; 5 incomplete. Check omitted-result counts before treating a report as exhaustive.

Terminal window
bee test summarize TestResults/results.trx --json

Read an existing TRX test report without loading the entire XML into the agent’s context. Bee reports totals, individual failures and bounded diagnostic details, together with the artifact’s hash. It does not run tests, read attachments or contact a service; initialization and MongoDB are not required. The initial reader accepts UTF-8 TRX with flat results; nested InnerResults are explicitly unsupported and produce an incomplete report.

Parsing completeness and test outcome are separate. A malformed, unfinished or unsupported report cannot establish success. Neither can a report with no tests. The test runner’s original process exit code is unknown: this command describes the selected artifact, not an independently verified run of the current checkout.

By default only non-passed result details are shown. Use --include-passed to inspect successful executions too. --limit caps visible records and --budget caps output characters; omitted and excluded-passed counts remain separate.

Results with the same name retain their separate execution identities. Long diagnostics and omitted results are marked explicitly. Use those markers to decide whether to inspect the original report; the summary does not reproduce all XML or captured standard output. Run bee test summarize --help for limits and exit-code details.

Use these tools when an agent needs evidence without reading every source file, log or report. The offline examples below use synthetic data. budget --by-item still requires your configured store; it is a forecast, not measured provider usage.

Output Bound and unit What to check
graph outline --compact --budget N Default 12,000; 2,048–65,536 UTF-16 code units, excluding final newline Column rows preserve declaration fields; counts, omittedMetadata, freshness and coverage
Graph relationship --output-budget-chars N At least 256 UTF-16 code units, including final newline; requires --json output.omittedItems, omittedByPath, legacyComplete and query truncation
test summarize --budget N Default 16,000; 2,048–65,536 UTF-16 code units, excluding final newline Returned/omitted records, excluded passed results and shortened diagnostics
code runs list/show --budget N Default 16,000; 2,048–65,536 UTF-16 code units, including final newline Integrity, original analysis outcome and omitted entries

UTF-16 units are not UTF-8 bytes or provider tokens: A uses one unit; many emoji use two. The legacy graph query --budget instead uses nominal tokens at three characters each and limits the query result, not the complete JSON envelope. Do not confuse it with outline’s character budget. A graph whole-output budget that cannot fit required metadata returns exit 4; increase it rather than accepting a stripped answer.

Compact outlines share result.columns and result.idPrefix; full ID = prefix + row suffix. This preserves declaration fields while removing repeated structure. Relationship-query --compact instead drops descriptive list fields. Neither guarantees lower billing or faster queries. Follow the coding example to select source lines from a real outline.

A graph build persists locally; relationship queries may build/refresh unless you request --no-refresh. Outline never refreshes. Saved compiler analysis is opt-in:

Terminal window
bee code analyze --engine roslyn --save-run --json
bee code runs list --json --limit 10 --budget 4096
bee code runs show <run-id> --json --limit 20 --budget 8192

Use the returned runId. show exit 0 means the saved artifact was read successfully, even if analysisExitCode is 1 or 5. integrity: verified checks stored bytes, not whether today’s source/toolchain/settings match. Reanalyze after relevant changes; see analysis history. Reading a saved report avoids rerunning analysis but does not establish provider prompt-cache use.

Use this when you have a supported local Codex rollout and its exact native thread UUID. No initialization or service is needed. A synthetic two-line rollout.jsonl illustrates the supported counter shape:

{"timestamp":"2026-09-22T12:00:00Z","type":"session_meta","payload":{"id":"11111111-1111-4111-8111-111111111111","cli_version":"0.155.0","originator":"codex_exec"}}
{"timestamp":"2026-09-22T12:00:01Z","type":"event_msg","payload":{"type":"token_count","info":{"total_token_usage":{"input_tokens":100,"cached_input_tokens":60,"output_tokens":20,"reasoning_output_tokens":5,"total_tokens":120},"last_token_usage":{"input_tokens":100,"cached_input_tokens":60,"output_tokens":20,"reasoning_output_tokens":5,"total_tokens":120},"model_context_window":258400}}}
Terminal window
bee budget --measured --harness codex --session 11111111-1111-4111-8111-111111111111 --log rollout.jsonl --json

This fixture returns exit 0 and schema 2: history.latest.cumulative contains input 100, cached input 60, uncached input 40, output 20 and total 120. Reasoning 5 is a subset of output; cached input is already included in input. The response stream is empty. cost, requestCount, usage and lastRequestUsage are null. These fixture numbers are not a savings measurement.

Keep response and history separate; never add streams, snapshots or segments. context_estimate, context_window_synthesis and history_with_estimated_baseline explicitly label estimated history bases. Counter consistency does not prove billing, complete request capture, model attention or that Bee caused a saving. complete describes the supplied snapshot; session finality remains unknown and child threads are not included automatically.

All three selectors are mandatory. Only supported Codex log shapes are accepted; other harnesses exit 5. Wrong thread identity, changing/corrupt input and unsupported records are incomplete (5); bad flags are invalid (4). Use a physical regular-file path: symlinks and ancestors that are symlinks are refused. Limits include 32 MiB/file, 4 MiB/line, 100,000 lines and 10,000 usage observations. A cut-off tail can retain earlier observations while still exiting 5.

Save this synthetic UTF-8 report as failure.trx. It describes a failure; it did not come from an executed test suite:

<TestRun id="33333333-3333-4333-8333-333333333333" xmlns="http://microsoft.com/schemas/VisualStudio/TeamTest/2010">
<Results><UnitTestResult testId="11111111-1111-4111-8111-111111111111" executionId="22222222-2222-4222-8222-222222222222" testName="PriceTests.Total" outcome="Failed"><Output><ErrorInfo><Message>Expected 42; got 41.</Message></ErrorInfo></Output></UnitTestResult></Results>
<ResultSummary outcome="Failed"><Counters total="1" executed="1" passed="0" failed="1" error="0" timeout="0" aborted="0" inconclusive="0" passedButRunAborted="0" notRunnable="0" notExecuted="0" disconnected="0" warning="0" completed="0" inProgress="0" pending="0"/></ResultSummary>
</TestRun>
Terminal window
bee test summarize failure.trx --limit 5 --budget 8000 --json

The result has parserStatus: complete, testVerdict: failed, one returned result and exit 1. The diagnostic is “Expected 42; got 41.” producerExitCode stays null. Removing the run ID makes the shape unsupported and returns incomplete/5. A summary hash identifies the bytes supplied, not their origin or currentness.

For duplicate prose, the context audit example above returns exit 1 when complete duplicate candidates exist; inspect totalFindings, returnedFindings and omittedFindings. Review whether both locations need the instruction before removing anything. Continue with rules and behavior to reduce irrelevant delivery rather than merely shortening an answer.

Use this when a test run has already produced a coverage artifact and you want to select the next source lines to inspect. It requires Bee 1.0.0-beta.14 or later and one existing Cobertura, OpenCover or LCOV report. It works offline without initialization, a database or model provider. For a small example, save this synthetic report as coverage.lcov; src/Calculator.cs is a label inside the report and need not exist to parse it.

SF:src/Calculator.cs
DA:10,3
DA:11,0
LF:2
LH:1
BRDA:10,0,0,3
BRDA:10,0,1,0
BRF:2
BRH:1
end_of_record
Terminal window
bee code coverage import coverage.lcov --format lcov --json

The result has status: complete, exitCode: 0 and an empty issues list. In files, src/Calculator.cs has lines and branches with state: known, total: 2, covered: 1 and percent: 50. These are separate measurements: lineData records three hits on line 10 and zero on line 11; branchData records one visited and one unvisited branch on line 10. No source was opened or tests run to produce this inspection.

Read status and issues before using any observations. A missing measurement is unknown, with null percentage; an explicitly zero denominator is not-applicable, also with null percentage. Neither means 100%. Exit 0 establishes parsing within the supported subset, not a coverage threshold or trusted QA acceptance. Exit 4 means invalid usage; exit 5 means incomplete, unsafe, unsupported, conflicting or over-limit input. Malformed or unsafe input may return only an error envelope. If observations remain in an incomplete report, do not treat them as complete coverage.

The input must be a regular UTF-8 file of at most 8 MiB, with no symlink anywhere in its path. Select the matching --format and read error or issues when parsing fails. Use a report exported in the supported subset; do not remove contradictory records merely to obtain exit 0. For example, XML with a DTD is unsupported. Bee writes its result to stdout; import does not save history, combine runs, read source or execute tests. Read, parse and output caps apply.

Keep rawSha256 with the original report to identify its bytes. Options such as --source-revision and --source-snapshot are unverified declarations: sourceBinding stays not-established, checkedSource stays not-read and qaAcceptance stays not-assessed. This does not establish current-source freshness or new-code coverage.

For an agent’s next step, first establish which source revision produced the report. Then inspect the reported file and line 11 for the unvisited line, and line 10 for its branch behavior. Use the coding workflow to select relevant declarations and callers; do not equate coverage with sufficient assertions or all repository files. Keep test outcomes and the runner’s original exit separately with test summaries. The beta.14 notes show all three format commands and the full declaration options.