Skip to content

Token preflight before a task

Requires Bee 1.0.0-beta.16 or later. The examples require a build containing bee token estimate. Check its help with bee token estimate --help.

Before starting a task, select the instructions, prompt and source excerpts you already know you will need. bee token estimate reads only the files and redirected stdin you explicitly select. It runs offline without initialization, a database, a provider call or a runtime encoding download. It does not execute the task, discover its context or predict the work an agent will do.

  1. Select known context, including required instructions. For a code review, this might be the review request, relevant source and test excerpts. Prepare alternative selections explicitly.
  2. Estimate each selection using the same encoding and assumptions. Inspect both the count and whether the selection still supports the task.
  3. Choose the context and perform the task separately. A lower count alone proves neither better quality nor token savings; removing necessary evidence can cause extra work.
  4. Afterwards, inspect an explicitly selected supported session log with the existing measured-usage command. Keep recorded observations separate from the preflight allocation.

The following Bash/zsh preparation creates synthetic text and an empty private Bee profile outside your checkout. pwd -P resolves the physical directory; on macOS /tmp itself is a symlink, so a root under /private/tmp is needed.

Terminal window
demo_dir="$(mktemp -d "${TMPDIR:-/tmp}/bee-token-demo.XXXXXX")"
demo_dir="$(cd "$demo_dir" && pwd -P)"
mkdir "$demo_dir/profile"
printf 'Hello world!\n' > "$demo_dir/prompt.txt"
printf 'Review the tests.\n' > "$demo_dir/context.txt"
chmod 444 "$demo_dir/prompt.txt" "$demo_dir/context.txt"
Terminal window
BEE_HOME="$demo_dir/profile" bee token estimate --root "$demo_dir" --file prompt.txt --json

The file has 13 UTF-8 bytes and 13 UTF-16 units, including the newline. The default utf16-ceil-div4@1 heuristic gives input.estimatedTokensPerCall: 4. It rounds utf16Units / 4 up for each input, then sums. It is an estimate, not a tokenizer count or an upper bound. Two one-unit inputs therefore contribute two, even though joining them would contribute one.

2. Select redirected stdin and an encoding

Section titled “2. Select redirected stdin and an encoding”
Terminal window
printf 'Hello world!' | BEE_HOME="$demo_dir/profile" bee token estimate --stdin --encoding o200k_base --json

This text has no trailing newline: 12 bytes, 12 UTF-16 units and 3 tokens under o200k_base. Stdin is read only with --stdin and must be redirected. Stdin-only input does not open a root directory. You may combine one --stdin with repeated --file selections.

Terminal window
BEE_HOME="$demo_dir/profile" bee token estimate --root "$demo_dir" \
--file prompt.txt --file context.txt --encoding cl100k_base \
--calls 3 --output-budget 2000 --json

The files contribute 3 and 4 tokens: 7 per call. --calls 3 means three independent calls with the same selected input each time. It does not simulate conversation growth, retries, tools or compaction. --output-budget 2000 is your allocation per call, including any intended reasoning allocation; it neither forecasts output nor enforces a generation limit.

Scenario field Value
inputEstimatedTokens 21 = 3 × 7
outputAllocatedTokens 6000 = 3 × 2000
totalAllocatedTokens 6021 = 3 × (7 + 2000)

4. Hide rows while preserving the aggregate

Section titled “4. Hide rows while preserving the aggregate”
Terminal window
BEE_HOME="$demo_dir/profile" bee token estimate --root "$demo_dir" \
--file prompt.txt --file context.txt --encoding cl100k_base \
--calls 3 --output-budget 2000 --limit 0 --json

contributors is empty, but the input and scenario totals above are unchanged. projection reports returned: 0, omitted: 2, omittedEstimatedTokens: 7. --limit controls visible contributor rows, not input selection or counting. Use the separated form --limit 0; --limit=0 is not accepted.

Terminal window
BEE_HOME="$demo_dir/profile" bee token estimate --root "$demo_dir" \
--file prompt.txt --encoding cl100k_base \
--calls 3 --output-budget 2000 --json

With the same encoding, calls and output allocation, selecting only prompt.txt gives 3 tokens per call, 9 scenario input tokens and a total allocation of 6009. This comparison measures a changed selection: it drops the instruction to review tests. Decide whether that omission is acceptable for the actual task. The smaller number is not evidence of equivalent work or realized savings.

6. Distinguish an explicit zero from unknown

Section titled “6. Distinguish an explicit zero from unknown”
Terminal window
printf '' | BEE_HOME="$demo_dir/profile" bee token estimate --stdin --output-budget 0 --json

Empty text is valid: its input count is zero, and this explicit zero output allocation gives total allocation zero. In examples 1 and 2, the omitted --output-budget leaves outputBudgetPerCall, outputAllocatedTokens and totalAllocatedTokens as null. Unknown is not zero. Even an explicit zero allocation does not establish what a later task will consume.

--encoding accepts heuristic (default), o200k_base or cl100k_base. The explicit encodings use packaged Microsoft.ML.Tokenizers and matching encoding data, version 2.0.0. They count ordinary text exactly under that selected implementation. Special-looking strings such as <|endoftext|> remain literal text. Missing encoding data fails with encoding_unavailable; it does not silently switch to the heuristic.

There is no default model mapping or provider compatibility claim. An exact selected-text count is not a full request count: messages, hidden framing and unselected system/tool definitions are outside it. Each input is counted separately and the counts are summed. If you need to count assembled text with separators, supply that complete text as a single file or stdin input.

Every result is one compact JSON document followed by a newline, even without --json. Successful estimates have schemaVersion: 1, status: complete and scope: selected_text_only. estimator identifies the method and versions; input includes exact byte/unit sizes and the per-call token count.

Contributor ordinals follow your original file/stdin option order, starting at 1. Rows are sorted by tokens descending, then UTF-16 units descending, then ordinal. projection accounts for omitted rows. Default --limit is 20; valid values are 0–64. --calls defaults to 1 and accepts 1–1000000; optional --output-budget accepts 0–1000000000. Numeric values use unsigned ASCII decimal digits.

The unknown list retains actual future output and reasoning, future tool results and turns, conversation growth and compaction, unselected system/tools/framing, cache behavior, and provider usage/billing. money is null: no pricing calculation is made. No cached-token or actual-provider-usage figure can be inferred from these counts.

File selections use relative / paths beneath a physical absolute --root; if omitted, the current directory is used. There is no directory scan. Traversal, links in the root or file path, hardlinks, duplicate physical files, and nonregular files such as FIFOs are refused. Select a regular copy inside the permitted physical root; do not use a symlink to bypass the boundary. Common sensitive filenames are refused, but this is not a general secret detector.

Input must be plain, strictly valid UTF-8. UTF-8 BOM is counted as U+FEFF; CRLF and trailing newlines are preserved. There is no trimming, normalization or replacement decoding. Malformed UTF-8, UTF-16/32 BOMs and NUL text are rejected. The report emits no content, filenames, paths or private input hashes, and the command does not persist the selection. File rechecks are best effort: snapshot: checked_unchanged_not_atomic is not an atomic snapshot guarantee.

Limit Behavior
64 selected inputs, including stdin Explicit selections only
256 KiB (262144 bytes) per input; 1 MiB (1048576 bytes) total Over-limit input is refused, not truncated
32 KiB (32768 bytes) JSON, including final newline Oversize output returns output_limit, not sliced JSON
Exact encoding work Across all inputs, the sum of squared UTF-8 byte lengths of natural pre-token pieces must be ≤67108864; a single piece can be at most 8192 bytes

Long unbroken text can exceed the exact-mode work cap even within the raw byte limits. The whole estimate then fails with tokenization_limit and null totals; there is no splitting, truncation or fallback. The encodings also have a 250 ms regex-operation timeout (encoding_timeout). The 10-second acceptance deadline is cooperative; blocked synchronous filesystem or output operations cannot promise a hard ten-second wall-clock exit.

Exit Meaning and next action
0 Complete estimate or help; still inspect scope, unknowns and omissions
4 Invalid arguments (invalid_usage); check spellings, ranges, selection and duplicate flags
5 Refused input, runtime or output failure; inspect errorCode and optional errorInputOrdinal

For example, stdin_not_redirected requires a pipe or redirection, missing requires an existing selected file, invalid_unicode/invalid_text require valid plain UTF-8, and file_byte_limit/total_byte_limit require a smaller explicit selection. Path/link refusals mean the selected path does not satisfy the physical-file boundary. An input/runtime failure has status: incomplete; argument errors have status: invalid. Both clear estimator, input, projection and scenario to null and return no contributors. No partial input total is published. An output-write failure can leave an incomplete transport prefix; exit 5 is not a usable JSON success.

Perform the task, then inspect recorded usage

Section titled “Perform the task, then inspect recorded usage”

Once the task has actually run, use a supported local log and its exact native thread UUID with the existing command:

Terminal window
bee budget --measured --harness codex \
--session 11111111-1111-4111-8111-111111111111 \
--log /absolute/physical/path/to/rollout.jsonl --json

Replace the example UUID and path with the exact thread and log you intend to inspect. All three selectors are required; Bee does not find the latest session automatically. The context size guide provides a synthetic log example and explains recorded counters, estimated history bases and incomplete evidence. Keep response and history streams separate; do not add them or infer billing, complete capture, cache savings or causation from preflight comparisons.

For context inspection and duplicate candidates, continue with context size; for syntax discovery, use the command reference.