Skip to content

Assess supplied QA evidence

Requires Bee 1.0.0-beta.15 with bee qa assess.

Give an agent one assessment of evidence you already have, then follow its named observations and next checks instead of reopening every raw report. The command reads local files and uses read-only Git. It does not run tests, report generators or attachments, authenticate a producer, grant acceptance or measure savings.

qa assess above only diagnoses supplied evidence. For a current checkout, bee qa run executes an explicit JSON task specification with criteria and checks; it does not infer commands from prose. Start with its live specification template and review every executable and argument before running it:

Terminal window
bee qa run --help --json
bee qa run --root /absolute/checkout --spec /absolute/task.json --json

The specification uses schemaVersion: 1, profile: "change-task", 1–16 criteria and 1–16 checks. A criterion names source paths and the checkIds that cover them. A repo-command check declares an absolute executable, an argument array, source paths, additional inputs and a 1–300 second timeout. code-analyze accepts native source/compiled or Roslyn compiled scope; it does not imply test coverage. Check inputs and source identity are captured before and after execution. The command may write files and start processes; it is not a sandbox. Ambient PATH is not forwarded, so use an absolute interpreter or a reviewed wrapper.

Exit 0 means verifiedCompletion=true for the declared checks only; exit 1 means repair_required, 4 invalid usage/spec, and 5 incomplete evidence or judgment needed. The report’s trust is in_process_only: saved JSON is not accepted as proof of a later run. Review raw outputs, check quality and subjective criteria. See advanced CLI workflows for a complete task example, and Plans for the separate owner-reported state.

Start in the checkout root. Choose a locally available base commit and either the current committed HEAD or worktree for current files. The example fixes the base to today’s HEAD and assesses the worktree; replace QA_BASE with the intended comparison commit when needed. A commit-valued head must match checkout HEAD and unchanged inputs. No commits are fetched.

The bundled bee-cli profile describes QA CLI integration for linux-x64. It is not a general profile for an arbitrary project, another platform or the whole release. Running the importer on macOS does not change that target.

Use a shell with Git and Python 3. This creates a new temporary bundle outside the repository and resolves its physical path, including /private/tmp on macOS:

Terminal window
export QA_ROOT="$(pwd -P)"
export QA_BASE="$(git -C "$QA_ROOT" rev-parse HEAD)"
export QA_HEAD=worktree
export QA_BUNDLE="$(python3 -c 'import pathlib,tempfile; print(pathlib.Path(tempfile.mkdtemp(prefix="bee-qa-demo-")).resolve())')"
bee qa assess --help

Keep the same checkout and bundle for the following steps. --root, --base, --head and --profile are required; the command needs no Bee initialization or service connection.

Terminal window
qa_exit=0
bee qa assess --root "$QA_ROOT" --base "$QA_BASE" --head "$QA_HEAD" \
--profile bee-cli --limit 5 --budget 16000 --json \
> "$QA_BUNDLE/preview.json" || qa_exit=$?
test "$qa_exit" -eq 5

Continue only if the exit check succeeds. Help exits 0, invalid/unknown/duplicate options exit 4, and every assessment deliberately exits 5. Exit 5 means the diagnostic assessment is not acceptance; it is not an observed test-process exit or proof that tests failed.

Without --evidence, expect decision: collect_evidence, missing_evidence and missing scope observations. Inspect preview.json before preparing real evidence. An unresolved source or an exhausted output budget is not a usable source identity.

This example writes synthetic data, not a test result produced by execution. It creates one TRX failure outside the profile, with a consistent test identity and counters. The schema-1 index declares the exact source manifest and policy from the preview, plus the artifact’s actual byte length and SHA-256. It leaves completion time and unavailable build/settings/toolchain/environment identities unknown.

Terminal window
python3 - <<'PY'
import hashlib, json, os
from pathlib import Path
bundle = Path(os.environ["QA_BUNDLE"])
preview = json.loads((bundle / "preview.json").read_text(encoding="utf-8"))
source = preview["source"]["manifestSha256"]
policy = preview["policyFingerprint"]
assert source and policy
trx = b'''<TestRun xmlns="http://microsoft.com/schemas/VisualStudio/TeamTest/2010"
id="11111111-1111-4111-8111-111111111111">
<TestDefinitions><UnitTest id="22222222-2222-4222-8222-222222222222">
<TestMethod className="Docs.Synthetic" name="Example"/>
</UnitTest></TestDefinitions>
<Results><UnitTestResult testId="22222222-2222-4222-8222-222222222222"
executionId="33333333-3333-4333-8333-333333333333"
testName="Synthetic example" outcome="Failed"/></Results>
<ResultSummary outcome="Failed"><Counters total="1" executed="1"
passed="0" failed="1" notExecuted="0"/></ResultSummary>
</TestRun>
'''
(bundle / "report.trx").write_bytes(trx)
index = {
"schemaVersion": 1, "runId": "docs-synthetic", "attemptId": "attempt-1",
"completedAt": None,
"bindings": {"source": source, "profile": "bee-cli", "policy": policy},
"acceptanceFingerprint": policy,
"artifacts": [{"kind": "trx", "relativePath": "report.trx",
"sha256": hashlib.sha256(trx).hexdigest(), "bytes": len(trx)}]
}
(bundle / "index.json").write_text(json.dumps(index) + "\n", encoding="utf-8")
print(bundle / "index.json")
PY

runId and attemptId label this supplied bundle. relativePath is relative to the index directory. acceptanceFingerprint is a policy binding field; setting it does not grant acceptance. Copying a source hash establishes only a content claim, and calculating an artifact hash does not authenticate its producer. Do not relabel an old real report with a new source hash to hide a mismatch.

Terminal window
qa_exit=0
bee qa assess --root "$QA_ROOT" --base "$QA_BASE" --head "$QA_HEAD" \
--profile bee-cli --evidence "$QA_BUNDLE/index.json" \
--limit 5 --budget 16000 --json > "$QA_BUNDLE/assessment.json" || qa_exit=$?
test "$qa_exit" -eq 5
python3 -m json.tool "$QA_BUNDLE/assessment.json"

With unchanged source, this example yields sourceAssociation: content_matches_claim, observedHistogram: {"Failed": 1} and decision: collect_evidence. The artifact has verification: bytes_checked; its parser can be complete while testVerdict is failed. Required profile scope is still missing. historical_failure retains the synthetic observation while unknown completion and other bindings prevent it from becoming a current-source repair conclusion.

Read decision, reasonCounts, scopeCounts, artifacts and next together. details names an artifact digest and occurrence when available, so an agent can inspect the relevant raw observation next. needs_judgment requests review; it does not classify a defect automatically. missing_result and not_executed do not become passes. A partial report can retain failures; malformed input remains inconclusive and cannot establish a validated test failure.

All assessments keep mode: diagnostics, trusted: false, acceptance: false and intentJudgment: required, including bundles that report only passes. observedProducerExit remains null. An optional producer record’s reportedExit is a supplied claim; a parser’s summarizerExitCode is another distinct value.

Terminal window
qa_exit=0
bee qa assess --root "$QA_ROOT" --base "$QA_BASE" --head "$QA_HEAD" \
--profile bee-cli --evidence "$QA_BUNDLE/index.json" \
--limit 200 --budget 65536 --json > "$QA_BUNDLE/expanded.json" || qa_exit=$?
test "$qa_exit" -eq 5

--limit accepts 0–200 details; --budget accepts 2048–65536 UTF-16 code units, excluding the final newline. They bound the projection after assessment. Check counts.returned, counts.omitted and counts.textTruncated; a shorter result does not mean fewer obligations. The example’s expanded result reveals the omitted details without changing its decision.

For output_budget_exceeded, repeat with a larger budget and the same evidenceIndexSha256; verify the source manifest and raw artifact digests too. The fallback may omit reason types explicitly. If source or evidence changed, treat it as a new assessment, not continuation of the old one. This is repeated inspection, not a result cache.

Keep the index and artifacts outside the repository, as regular files with no symlink or reparse-point component in their paths. Use portable relative artifact paths, without .., absolute paths or backslashes. Inputs must be strict UTF-8. The index allows at most 1 MiB and 64 artifacts; each artifact at most 16 MiB, together at most 64 MiB. Source capture and parsers have additional limits, and the command has a ten-second overall time boundary. Over-limit, inaccessible or changing input remains explicit.

source_mismatch means the declared source does not match captured content. stale_evidence means the declared completion is over 24 hours old; absent completion remains unknown and future completion needs judgment. Older or mismatched failure observations remain visible without proving a current defect. No supported option overrides these limits to trust or accept a bundle. This command supplies no test runner, test selection, mutation testing or additional backend, Vue or mobile profiles.

Command reference · Beta.15 notes