Reference

Command-line reference

Every command, flag, exit code, environment variable, policy key and risk signal of the current Probe (latest release: v0.5.1). Entries marked v0.4 make a binary older than v0.4.0 exit 3; see release ordering before pinning one in CI. Press / to search.

This page covers Probe for code, the command-line tool. Watching Word, Excel or PowerPoint files? See Probe Desktop.

Synopsis

probe init [--repo PATH] [--language go|typescript|javascript|python|rust]
probe lint [--base main] [--head HEAD] [--ci] [flags]
probe review [--base main] [--reviewer=false] [--ci] [flags]
probe review [flags] BASE..HEAD
probe plan --intent-file FILE [--base main] [--ci]
probe review --plan .probe/PLAN.json [flags]
probe report [--input .probe/confidence-report.json] [--out DIR] [--format LIST] [--report-url URL]
probe version
  • lint analyzes a change statically. It never executes repository code and never calls an AI provider.
  • review adds sandboxed checks in Docker, the opt-in evidence stages (changed baseline tests, impacted tests, differential fuzzing, mutation, dependency preparation) and, when a model is configured, an AI investigation.
  • plan asks the configured provider for an implementation plan before any code is written, evaluates it with fixed rules, and turns it into a contract that lint --plan and review --plan check the diff against.
  • Only committed files are analyzed; uncommitted edits and untracked files are ignored.
  • Reports are written to .probe/ in the repository by default. Console output stays short.
  • Flags accept one or two dashes (-ci or --ci). Booleans are set with --ci or --ci=true and cleared with --checks=false. A value is given as --base main or --base=main.
  • Run probe <command> --help for the flags of a command.

Choosing revisions

Both sides of a comparison resolve to immutable commit IDs, recorded in the report. Renames, deletions, binary files and file type changes are handled.

FormCompares
--base main (default)From the merge base of main and --head to --head, like a pull request. Commits added to main since the branch point are not part of the change.
--base main --exactFrom the tip of main directly to --head.
BASE..HEADThe two exact commits, for example HEAD~1..HEAD to review the last commit. Overrides --base, --head and --exact.
BASE...HEADFrom the merge base of the two revisions to HEAD.

A range may appear before or after the flags: probe review main..HEAD --ci and probe review --ci main..HEAD are equivalent. At most one range is accepted. Any revision Git understands works: branches, tags, origin/main, HEAD~3 or a commit ID. CI jobs need enough history to resolve the base (for example fetch-depth: 0).

The trusted policy is always read from the base branch, never from the change under review, so a pull request cannot loosen its own rules.

probe lint

Static analysis of a change: Git comparison, risk signals and a report. No Docker, no provider call, no execution of repository code. Safe to run on untrusted branches.

FlagDefaultDescription
--base REVmainBase branch or revision. The comparison starts at its merge base with --head; its tip supplies the trusted policy.
--head REVHEADCandidate revision under review.
--exactfalseCompare the base revision itself instead of the merge base.
BASE..HEAD—Positional range; see Choosing revisions.
--repo PATH.Repository directory. Any directory inside the work tree works.
--config PATHpolicy of the baseUse this local policy file instead of .probe.json from the base branch. Only pass a file you trust.
--out DIR.probeReport directory, relative to the repository root. The path may not contain symbolic links.
--format LISTmarkdown,jsonComma-separated report formats to write: markdown, json, sarif v0.4, pr-comment v0.4. See Output files.
--report-url URL v0.4—An https link to the full report, cited in PR_COMMENT.md. Requires the pr-comment format; a URL that looks like it carries a credential is refused with exit 3.
--impacttrueBuild the static impact index of changed functions. --impact=false skips it.
--plan PATH—A PLAN.json written by probe plan: check the diff against its contract and add a plan_drift section and signals.
--cifalseExit with code 2 when human review is required. See Exit codes.
--intent TEXT—Pull request intent or acceptance criteria, recorded in the report (at most 64 KiB).
--intent-file PATH—Read the intent from a UTF-8 file. Cannot be combined with --intent.
--checks, --reviewerfalseMust stay false: lint exits with code 3 if either is enabled. Use review. The review-only stage flags (--base-tests, --impacted-tests, --fuzz, --cache-dir, --parallel, --deadline, --allow-prepare-network) are refused the same way.
probe lint HEAD~1..HEAD                 # the last commit
probe lint --base origin/main --ci      # this branch as a pull request, for CI
probe lint --base main --format json    # JSON report only
probe lint --base main --plan .probe/PLAN.json   # scope drift against an approved plan

probe review

Everything lint does, plus the policy's test, typecheck, build and coverage commands in disposable Docker containers, plus an AI investigation when a model is configured. Requires Docker with Linux containers and a preloaded sandbox.image.

FlagDefaultDescription
--base REVmainBase branch or revision. The comparison starts at its merge base with --head; its tip supplies the trusted policy.
--head REVHEADCandidate revision under review.
--exactfalseCompare the base revision itself instead of the merge base.
BASE..HEAD—Positional range; see Choosing revisions.
--repo PATH.Repository directory.
--config PATHpolicy of the baseUse this local policy file instead of .probe.json from the base branch, for example to try a policy before committing it. Only pass a file you trust: it decides what runs and where source is sent.
--out DIR.probeReport directory, relative to the repository root. Artifacts go to DIR/artifacts/.
--format LISTmarkdown,jsonComma-separated report formats to write: markdown, json, sarif v0.4, pr-comment v0.4.
--report-url URL v0.4—An https link to the full report, cited in PR_COMMENT.md. Requires the pr-comment format.
--cifalseExit with code 2 when human review is required, including high-risk signals, unverified areas or incomplete checks.
--checkstrueRun the configured checks in the sandbox. --checks=false skips them; a configured reviewer may still run sandbox experiments.
--reviewerautoAI investigation. Enabled automatically when a model is set in the trusted policy or PROBE_REVIEWER_MODEL. --reviewer=false disables every provider call; --reviewer fails with exit 3 if no model is configured.
--max-iterations Npolicy (20)Override reviewer.max_iterations for this run, 1 to 100.
--intent TEXT—Pull request intent or acceptance criteria (at most 64 KiB). Recorded in the report and given to the reviewer.
--intent-file PATH—Read the intent from a UTF-8 file. Cannot be combined with --intent.
--allow-networkfalseGive sandbox containers network access. Takes effect only if the trusted policy also sets sandbox.network: true.
--no-networkfalseForce sandbox networking off. Does not affect reviewer API calls; add --reviewer=false for that.
--plan PATH—Check the diff against the contract of a PLAN.json. A drifted change requests human review (exit 2 with --ci), never exit 1.
--impacttrueBuild the static impact index of changed functions. --impact=false skips it (and cannot be combined with --impacted-tests).
--impacted-tests v0.4falseRun the unchanged Go tests that the impact index lists as reaching a changed function, on baseline and candidate. A test passing on the base and failing on the candidate is FAILS_ON_CANDIDATE: a reason to review, not a defect.
--base-tests v0.4falseRun the baseline version of each Go test function the change modified or removed, on baseline and candidate code.
--fuzz v0.4trueRun the differential fuzzing that the policy's fuzz object configures. --fuzz=false records the stage as disabled.
--cache-dir DIR v0.4—Opt-in cache of baseline executions, outside the repository and the output directory. A replayed baseline run never supports a positive result.
--parallel N v0.41Run the initial checks N at a time, 1 to 4.
--deadline D v0.4—Overall time limit of the run, from 1m to 24h; 30 seconds of it are kept for cleanup and the report.
--allow-prepare-network v0.4falseGive the prepare container network access, only if the policy's prepare.network also enables it. Checks stay offline.
probe review --base main --ci                        # typical pull request run
probe review HEAD~1..HEAD --reviewer=false           # checks only, no provider
probe review --base main --checks=false --reviewer=false   # static only, like lint
probe review --base main --intent-file PR.md         # give acceptance criteria
probe review --base main --impacted-tests --base-tests --ci   # run tests on both revisions
probe review --base main --format markdown,json,sarif,pr-comment --report-url "$RUN_URL"

Dependency preparation runs first, when the policy has a prepare object. Checks then run in the order test, typecheck, build, then coverage, followed by the evidence stages and the reviewer. Each container is non-root, has a read-only root and source mount, no added capabilities, no network by default, and the CPU, memory, PID and time limits of the sandbox policy. The Docker socket, your working checkout and API keys are never mounted. If Docker or the image is missing, the run ends with exit 4; nothing ever falls back to your host.

probe plan

Before the change is written: the configured provider simulates, read-only, how it would implement an intent at the base commit and submits a structured plan (files, symbols, dependencies, steps). Probe evaluates it with fixed rules and writes PLAN.json and PLAN.md. The model produces the plan; it never judges its risk. Nothing is modified or executed.

FlagDefaultDescription
--intent-file PATH, --intent TEXT—The change to plan. Required; at most 64 KiB of UTF-8, read like the review intent.
--base REVmainThe revision the plan starts from. Its tip supplies the trusted policy.
--config PATHpolicy of the baseExplicit trusted local policy.
--out DIR.probeOutput directory, relative to the repository.
--cifalseExit 2 when a category is flagged or something was left unverified.
--max-iterations NpolicyOverride the provider iteration budget, 1 to 100.
--reviewertrueA provider is mandatory: without a configured model, or with --reviewer=false, plan exits 3 and writes nothing.

The planner's tools are read-only and work on a snapshot of the base commit: list_files, read_file, search_code and the index-backed find_references, inspect_symbol and find_callers, then submit_plan once. The assessment flags four categories from the plan and the base commit only:

CategorySignals
Critical partsplan_critical_path (a planned path matches sensitive_paths), plan_sensitive_symbol (an authentication, authorization or payment name).
Architectureplan_exported_signature, plan_dependency_change, plan_new_package.
Regression riskplan_wide_impact (10 or more callers), plan_untested_impact (callers and no reaching test), measured on the static index of the base commit.
Other major changeplan_file_deletion, plan_large_scope (20 or more files), plan_inconsistent, plan_unmeasured.

Exit codes: 0; 2 with --ci when a category is flagged; 3 usage or configuration; 4 provider failure or no accepted plan.

Scope drift: --plan

lint --plan and review --plan read the plan's contract and compare the real diff with it. Each difference in the diff also becomes a plan_drift signal. The section is drifted when an item is medium or higher.

ItemSeverityWhen
unplanned_filemediumA changed file the plan does not list (unplanned_test_file, low, for a test file).
unannounced_exported_changehighAn exported Go declaration changed or removed without being announced as signature or remove.
unannounced_critical_pathhighA changed path matches a critical glob of this review's policy and the plan did not list it.
unannounced_dependency_changemediumA dependency manifest changed without being declared.
planned_file_untouchedlowA planned file the change does not touch (section only).
probe plan --intent-file task.md --ci                     # 1. plan; exit 2: a human validates it
# 2. an agent implements the plan
probe review --base origin/main --plan .probe/PLAN.json --ci   # 3-4. exit 0: merge; exit 2: review

The plan gate

review --plan ends with a decision, plan_drift.decision, printed on stdout and at the top of the Plan Conformance section. With --ci, it is the exit code.

DecisionWhen
no_human_review_required (exit 0)All of: the plan, re-assessed by the review from its proposal at its base commit with this review's trusted policy, flags no category and has no measurement gap; the change conforms to the plan and to its base; at least one check ran and all passed; no reproduced issue, high or critical signal, unverified area or other section requesting review.
human_review_required (exit 2)Anything else. Every reason is listed in plan_drift.decision_reasons. lint --plan runs no check, so it always requires review.

The flags and contract stored in PLAN.json are never trusted: the contract is re-derived from the proposal, so an edited plan cannot widen the scope unassessed, and probe report recomputes the decision from the recorded report. The gate is a process decision, not a verdict on correctness. A plan that raises nothing is not proof that the change is safe, and drift does not check that the change implements the intent. See pre-change plans and the CI recipe.

probe init

Write a starter .probe.json for the project. It refuses to overwrite an existing file. Review the commands and image, then commit the file to your base branch.

FlagDefaultDescription
--repo PATH.Directory in which to create .probe.json.
--language NAMEdetectedgo, typescript, javascript, python, rust or unknown. Detected from go.mod/go.work, Cargo.toml, tsconfig.json, package.json, then pyproject.toml/setup.py/requirements.txt. Selects the defaults.

probe report

Re-render a saved JSON report, for example to regenerate Markdown or produce SARIF in CI. Nothing is re-run, and re-rendering does not authenticate the evidence it contains; every status is re-derived from the recorded checks, so an edited report is corrected rather than trusted.

FlagDefaultDescription
--input PATH.probe/confidence-report.jsonSaved JSON report (version 1, at most 64 MiB).
--out DIR.probeOutput directory, relative to the current directory.
--format LISTmarkdown,jsonComma-separated formats to write: markdown, json, sarif v0.4, pr-comment v0.4.
--report-url URL v0.4—An https link to the full report, cited in PR_COMMENT.md.

probe version and help

CommandPrints
probe version, --versionThe version, for example probe v0.5.1, and the license notice.
probe help, -h, --helpThe synopsis. With no argument, Probe prints the same text.
probe <command> --helpThe flags of a command, with their defaults.

Exit codes

CodeMeaning
0Report completed without a reproduced high or critical issue. Without --ci, unresolved areas do not change this code.
1A high or critical hypothesis is supported by a differential failure: a generated test passed on the base and failed on the candidate. Nothing else produces 1: not a fuzzing divergence, a surviving mutant, a failing impacted or baseline test, nor plan drift.
2With --ci only: human review required, including high-risk signals, unverified areas, incomplete checks, a FAILS_ON_CANDIDATE test, a fuzzing divergence or inconclusive stage, plan drift, a plan that raised a category (the plan gate), or (for plan) a flagged category.
3Invalid arguments, Git comparison, or trusted configuration (including an unknown policy key or command name).
4Operational error in the harness, analysis or report writing, such as Docker or the sandbox image being unavailable, a check that could not run, a failed dependency preparation, or a provider failure during plan.

No code means "approved". A 0 says that nothing was reproduced, not that the change is correct.

Environment variables

The AI provider belongs to the deployment rather than to the repository, so these settings can come from the environment. Nothing else can: the image, commands, budgets and sensitive paths are always decided by the trusted policy.

VariableEffect
PROBE_REVIEWER_ENDPOINTOverrides reviewer.endpoint. Same URL rules as the policy.
PROBE_REVIEWER_MODELOverrides reviewer.model, and on its own enables the reviewer during review.
PROBE_API_KEYAPI key. The name is set by reviewer.api_key_env; this is the default.
PROBE_API_KEY_FILEPath of a file containing the key, used when the variable above is unset. The file must be readable, or the run fails with exit 3.
/run/secrets/PROBE_API_KEYDocker secret read last, when neither variable is set. Absence is not an error.

A blank variable counts as unset. Key files are read whole with surrounding whitespace stripped, bounded to 8 KiB and must hold a single line. When the reviewer runs, the log names where each value came from, never the value itself. Whoever controls this environment chooses where redacted source is sent: keep it out of jobs that run untrusted fork code.

Output files

PathContents
.probe/CONFIDENCE_REPORT.mdThe report to read: change summary, automated checks, investigation summary, reproduced issues, unverified areas, suggested human review, review surface, changed-line execution, recorded evidence and artifacts.
.probe/confidence-report.jsonThe same data for tools: commit IDs, changed lines, signals, checks, hypotheses, evidence, audit events and artifact hashes. See the JSON schema.
.probe/confidence-report.sarif v0.4With --format sarif: a SARIF 2.1.0 log for code-scanning tools, holding only findings backed by recorded sandbox evidence. No finding is not approval.
.probe/PR_COMMENT.md v0.4With --format pr-comment: a pull-request comment with a status block and every evidence-backed finding.
.probe/PLAN.json, PLAN.mdWritten by probe plan: the model-written proposal, the deterministic assessment and the contract (schema).
.probe/artifacts/Check logs, coverage profiles, generated test sources, test results, mutant patches and fuzzing records, each referenced by its SHA-256 hash in the report.

Add .probe/ to .gitignore. The console prints one summary: files and lines changed, number of signals and reproduced issues, the focused review surface and changed-line execution.

Risk signals

Signals are reasons to look, not confirmed defects. Go files get syntax-aware analysis; TypeScript/JavaScript, Python, Rust and other languages use labelled text heuristics. Signals on the same lines are merged into one review range in the report. File-level signals (sensitive_path, dependency_change, migration_change, infrastructure_change, binary_change, file_deleted, file_type_change, large_change, branch_growth, no_test_change, prepare_input_changed, plan_drift) concern the file, not a line: the report lists them under the file as whole file, and the JSON marks them "scope": "file". Their line only places them on the file's first changed line.

KindSeverityRaised when
sensitive_pathhighA changed path matches a sensitive_paths glob.
private_keycriticalAn added line starts a PEM private key block (RSA, DSA, EC, OpenSSH, PGP, encrypted) or a PuTTY key. The key is never copied into the report.
hardcoded_secrethighAn added line holds a credential shape (AWS, GitHub, GitLab, Slack, Stripe, Google, LLM provider, npm, Twilio/SendGrid, Azure storage keys, JWTs) or assigns a quoted literal of 8 or more characters with a digit or symbol to a secret-named key (password, api_key, token, client_secret…). Placeholders, templates and environment or vault lookups are ignored; the value is masked in the evidence.
credential_in_urlhighAn added URL carries user:password@ or a secret-named query parameter (api_key=, token=, password=…).
tls_verification_disabledhighInsecureSkipVerify: true, verify=False, rejectUnauthorized: false, NODE_TLS_REJECT_UNAUTHORIZED=0, curl -k, sslmode=disable, TLS 1.0/1.1 minimums and similar.
excessive_permissionshighWorld-writable modes (chmod 777, 0666 in file APIs), privileged containers, privilege escalation, runAsUser: 0, USER root, host PID or network, SYS_ADMIN, a mounted Docker socket, NOPASSWD: ALL.
protection_disabledhighCSRF disabled or exempted, wildcard CORS origins, insecure cookies, template autoescaping off, JWT none algorithm or unverified signatures, SELinux, firewall, seccomp or AppArmor disabled, public buckets or 0.0.0.0/0 ingress, workflow permissions: write-all.
debug_enabledmediumDEBUG = True, debug: true, app.run(debug=True), FLASK_DEBUG=1, gin.DebugMode and similar.
hardcoded_email, hardcoded_iplowAn added e-mail address, or an IPv4 address inside a string or URL, outside tests and documentation. Documentation, loopback and example ranges are ignored.
auth_changehighA Go function whose name suggests authentication or authorization has its body changed, or a changed line mentions authorization, authentication, JWT, bcrypt, argon2, CSRF or CORS.
sensitive_function_changehighA Go function whose name suggests payments (payment, refund, charge, capture, withdraw, deposit, balance…) has its body changed.
validation_removedhighA removed line called a validation, assertion or sanitizing function, or raised a validation error.
public_api_changehigh mediumGo: an exported declaration was removed or changed (high) or added (medium). Other languages: a line looks like a public declaration (medium).
database_writehighA changed line looks like a database mutation or transaction (INSERT, UPDATE … SET, .Exec(, .Commit(…).
dynamic_executionhigheval, process execution, subprocess, child_process, innerHTML and similar.
type_suppressionhighType or lint suppression such as @ts-ignore, as any, unsafe, nolint, eslint-disable.
migration_changehighA path containing migration or ending in .sql changed.
infrastructure_changehighA GitHub workflow, Dockerfile or Terraform file changed.
file_type_changehighGit reports a type change, for example a file becoming a symbolic link.
network_changemediumA changed line makes HTTP, gRPC or socket calls, or contains a URL.
error_handling_changemediumError checks, wrapping, panic/recover, catch or except changed.
dependency_changemediumA dependency manifest or lockfile changed (go.mod, package.json, Cargo.lock, pyproject.toml…).
uncovered_changemedium lowAdded Go lines were not executed by the recorded coverage run. Low when the coverage run itself failed.
branch_growthmediumAt least five more lines with branch constructs were added than removed in a file. Not a complexity metric.
large_changemediumMore than 400 lines changed in one file.
file_deletedmediumA tracked file was removed.
binary_changemediumA binary file changed and cannot be analyzed as text.
analysis_limitedmediumGo declaration analysis could not complete for a file, for example because it does not parse, or the impact index was limited or unavailable for changed functions (symbol impact_index).
test_focus_added v0.4highA focus marker was added to a test file (it.only, fdescribe…).
test_assertion_removed, test_case_removed v0.4mediumA test file lost more assertion lines, or more test declarations, than it gained. Go, TS/JS, Python and Rust rules.
test_skip_added, test_expectation_relaxed v0.4mediumA skip marker was added (t.Skip, it.skip, @pytest.mark.skip, #[ignore]…), or exact expectations were replaced by looser ones in a hunk.
surviving_mutant v0.4mediumA mutation of an added Go line made no test of its package fail. The mutant may be equivalent.
prepare_input_changed v0.4mediumThe change touches a file listed in the policy's prepare.inputs; the dependency layer was still built from the base commit.
plan_drifthigh medium lowWith --plan: the diff leaves the plan's contract. See scope drift.
impacted_callerlowUnchanged code calls a changed function, according to the impact index. At most 10 per function and 100 per run.
no_test_changelowA source file changed and no test file in the same directory or with the same name stem changed. Existing coverage is not measured.
todo_addedlowA TODO, FIXME, HACK or XXX marker was added.

Impact analysis

With --impact (the default), lint and review build a static index of the candidate commit on the host, from committed Git objects, without running repository code. For each changed function it lists the places in unchanged code that call it and the existing tests that reach it within 3 calls, adds low impacted_caller signals, and backs the reviewer's find_references, inspect_symbol and find_callers tools.

LanguageHow it is indexedResolution
GoParsed and type-checked with the Go standard library, package by package, with linux/amd64 build constraints. Interface calls are possible dispatch.static, interface
TypeScript/JavaScript, Python, RustScanned lexically: functions, methods and tests are found from tokens, and a call is linked by name to the same-named declarations of the same language, preferring the caller's class, the qualifying module or type and the same file. node_modules, dist, target, venv and similar directories are skipped.name

Every answer is approximate: a listed caller is a place to review, and an absent caller is not proof that none exists. --impacted-tests runs reaching Go tests only. See impact analysis.

Languages

FeatureGoTS/JSPythonRust
Risk signals and test-weakening rulessyntax-awaretexttexttext
Impact index and reviewer symbol toolstype-checkedlexicallexicallexical
Sandbox checks and init defaultsyesyesyesyes
Verified generated testsyesyes (Jest/Vitest JSON)run, not verified by nameno default command
Differential fuzzingyesyes——
Coverage, mutation, --base-tests, --impacted-testsyes———

Policy file: .probe.json

The policy decides which commands run, in which image, with which limits, and whether an AI provider is used. It is read from .probe.json on the base branch, or from --config PATH. When there is none, built-in defaults for the detected language apply. Keys you omit take the default values shown below.

The file must be a single JSON object of at most 1 MiB. Unknown keys, unknown command names and duplicate keys are rejected with exit 3, so an older binary rejects a key it does not know.

KeyTypeDefaultDescription
versioninteger1Policy format version. Must be 1.
languagestringdetectedLanguage written by init. Informational.
fuzz, mutation, prepare v0.4objectabsentOpt-in evidence stages: differential fuzzing, mutation of added lines and trusted dependency preparation. init never writes them.
{
  "version": 1,
  "language": "go",
  "commands": {
    "test": ["go", "test", "./..."],
    "typecheck": ["go", "vet", "./..."],
    "build": ["go", "build", "./..."],
    "generated_test": ["go", "test", "{package}"],
    "coverage": ["go", "test", "-covermode=count", "-coverprofile={coverage_out}", "./..."]
  },
  "sandbox": { "image": "golang:1.26-bookworm", "network": false, "timeout_seconds": 120,
               "max_runtime_seconds": 600, "max_output_bytes": 65536, "memory_mb": 1024, "cpus": 2 },
  "reviewer": { "endpoint": "https://api.openai.com/v1/chat/completions", "model": "",
                "api_key_env": "PROBE_API_KEY", "max_iterations": 20, "max_generated_tests": 10,
                "timeout_seconds": 600, "max_input_bytes": 131072 },
  "sensitive_paths": ["**/auth/**", "**/payment*/**", "**/migrations/**", ".github/workflows/**", ".probe.json"]
}

commands

Each command is an argv array, not a shell string: ["npm", "test"], never "npm test". At most 128 arguments; configure only the checks your project provides.

KeyPlaceholdersDescription
commands.test—Test suite, run by review.
commands.typecheck—Type or static check, for example go vet or tsc --noEmit.
commands.build—Build command.
commands.generated_test{file}, {package}, {results_out}How the AI reviewer's temporary tests (and --impacted-tests) are run on base and candidate. Go's default uses the test's package so it can exercise unexported code. Verified Go experiments need a single standalone target placeholder; a Jest or Vitest command writing a JSON report to {results_out} makes TS/JS experiments verifiable by test name.
commands.coverage{coverage_out} (exactly once)Go only. Measures which added lines a sandbox run executed. Runs last, in addition to test, within sandbox.max_runtime_seconds. Add -coverpkg=./... to attribute execution across packages.

sandbox

KeyDefaultAllowedDescription
sandbox.imagegolang:1.26-bookwormimage namePreloaded image with the toolchain and dependencies. Never pulled by Probe; pin it by digest if you can.
sandbox.networkfalsebooleanAllow container networking. Also requires --allow-network on the command line.
sandbox.timeout_seconds1201–3600Time limit of each command.
sandbox.max_runtime_seconds6001–7200Total sandbox time for the run, shared by checks and experiments.
sandbox.max_output_bytes655361024–4194304Captured output kept per command.
sandbox.memory_mb1024128–32768Memory limit of each container, in MiB.
sandbox.cpus21–32CPU limit of each container.

reviewer

KeyDefaultAllowedDescription
reviewer.model""model IDTool-capable model. Empty disables the reviewer. Overridden by PROBE_REVIEWER_MODEL.
reviewer.endpointhttps://api.openai.com/v1/chat/completionsURLChat Completions endpoint with function calling; a /v1 base URL also works. HTTPS, or HTTP on loopback only. No credentials, query or fragment; redirects are refused.
reviewer.api_key_envPROBE_API_KEYvariable nameEnvironment variable holding the key; <NAME>_FILE and /run/secrets/<NAME> follow it. Empty for a provider that needs no key.
reviewer.max_iterations201–100Model turns per investigation. --max-iterations overrides it.
reviewer.max_generated_tests100–100Temporary tests the reviewer may create.
reviewer.timeout_seconds6001–1800Total time for the investigation.
reviewer.max_input_bytes1310724096–2097152Bound on source context sent to the provider.

sensitive_paths

Globs of paths that always raise a high sensitive_path signal when changed. Paths are relative to the repository root with forward slashes; * matches within a path segment and ** across segments. Absolute paths, backslashes and .. are rejected.

Default globCovers
**/auth/**Any auth directory.
**/payment*/**payment, payments and similar directories.
**/migrations/**Database migrations.
.github/workflows/**CI workflows.
.probe.jsonThe policy itself.

fuzz v0.4

Runs the same seeded inputs through each changed package-level Go function and changed exported TS/JS function (with an unchanged signature) on baseline and candidate, and compares what the two revisions recorded. A diverged function is an observation, never a defect: it requests review (exit 2 with --ci) and never produces exit 1. {} enables it with the defaults.

KeyDefaultAllowed
fuzz.max_functions81–32 functions per review
fuzz.max_packages41–16 Go packages and TS/JS modules
fuzz.max_inputs641–256 inputs per function
fuzz.call_timeout_ms100010–10000, at most the command timeout
fuzz.max_runtime_seconds2401–7200, within sandbox.max_runtime_seconds

mutation v0.4

Makes small deterministic changes to the added lines of changed non-test Go files and runs the package's tests once per mutant. A mutant no test notices becomes a medium surviving_mutant signal; no score is computed. All four keys are required.

"mutation": {
  "command": ["go", "test", "-json", "-count=1", "-failfast", "{package}"],
  "max_mutants": 20, "timeout_seconds": 60, "max_runtime_seconds": 300
}
KeyAllowed
mutation.commandgo test with exactly one -json and one standalone {package}; flags that change the selected tests or binary (-run, -exec, -o…) are refused.
mutation.max_mutants1–200.
mutation.timeout_seconds1 to sandbox.timeout_seconds, per run.
mutation.max_runtime_secondstimeout_seconds to sandbox.max_runtime_seconds, inside the shared budget.

prepare v0.4

Lets the trusted base-branch policy build the dependency layer: before any candidate code runs, its command runs once, in one bounded container, on dependency files exported from the base commit, and the container becomes the local image of every check. A later review of the same base and inputs reuses it. When it produces no image, the review exits 4; it never falls back to the unprepared image.

"prepare": { "command": ["go", "mod", "download"], "inputs": ["go.mod", "go.sum"], "network": true }
KeyDefaultDescription
prepare.commandrequiredArgv of 1 to 128 arguments; no placeholder is substituted.
prepare.inputsrequired1 to 64 repository-relative patterns (* never crosses /, no **).
prepare.networkfalseNetwork for the build container, only with --allow-prepare-network and without --no-network.
prepare.user"sandbox""sandbox" (the identity of checks) or "root".
prepare.timeout_seconds6001–3600.
prepare.envnoneUp to 32 variables baked into the derived image; names that would change what checks execute (GOFLAGS, NODE_OPTIONS, LD_PRELOAD, PROBE_*…) are refused.
prepare.max_added_mb40961–65536 MiB the derived image may add.

Language defaults

Written by probe init, and used when the base branch has no policy. Stock Node.js, Python and Rust images do not contain your dependencies: build an image that does, or add a prepare object, and adapt the commands.

LanguageImageCommands
gogolang:1.26-bookwormtest go test ./... · typecheck go vet ./... · build go build ./... · generated_test go test {package} · coverage go test -covermode=count -coverprofile={coverage_out} ./...
typescript, javascriptnode:22-bookwormtest npm test · build npm run build · generated_test npx --no vitest run {file} --reporter=json --outputFile={results_out}
pythonpython:3.13-bookwormtest python -m unittest discover · generated_test python -m unittest {file}
rustrust:1-bookwormtest cargo test --workspace --offline · typecheck cargo check --workspace --all-targets --offline · build cargo build --workspace --offline · no generated_test
unknowngolang:1.26-bookwormNo commands: review runs no checks until you add them.

AI reviewer

Optional. When a model is configured, review lets it investigate the change with bounded tools: reading redacted files and diffs, searching source, looking up references and callers in the static index, running existing checks, and creating and running temporary tests on base and candidate. It has no shell and no URL fetcher. Its claims stay hypotheses until a test reproduces a difference.

{
  "reviewer": {
    "endpoint": "https://your-provider.example/v1",
    "model": "your-tool-capable-model",
    "api_key_env": "PROBE_API_KEY"
  }
}
export PROBE_API_KEY=…                       # from your shell or CI secret store
probe review --base main --config .probe.json   # try it before committing
probe review --base main --reviewer=false            # turn it off for one run
  • Only the trusted policy, --config or the deployment environment can enable the reviewer; a pull request cannot.
  • Bounded, redacted source context is sent to the provider. Masking of secrets is best effort; use a local provider (for example http://127.0.0.1:1234/v1) if source must stay local.
  • API keys never enter test containers. lint never calls a provider.
  • Provider failures and exhausted budgets keep the deterministic results and mark the investigation incomplete; with --ci that requires human review.