Open source · For code and for office documents

Review what matters.

Changes pile up faster than anyone can read them: pull requests written with AI assistants, budgets and contracts edited by many hands. Probe compares each new version with the one you last trusted, puts the few modifications that need your judgment in front of you, and shows why.

Latest v0.5.1 · Findings come from fixed rules, not from a model · AI optional.

A person working calmly at a desk inside a bright circle, at the center of a tunnel of tangled cables, notifications and dashboards
19 / 125 lines need you Probe filters the rest.
Two tools, one discipline

The same review method, for developers and for business teams.

Both tools start from the version you last reviewed, flag the risky modifications with fixed, explainable rules, and keep a record of what was checked. An AI model can help explain a finding; it never decides what is flagged, and nothing leaves your machine unless you choose it.

Probe for code

Review AI-written pull requests

For developers, tech leads and CI pipelines.

  • Maps a Git diff to risk signals: authentication, payments, removed validation, public APIs, dependencies.
  • Runs your tests in disposable containers and tries to reproduce issues on both revisions.
  • Checks an agent's plan before the code exists, then lets conforming low-risk changes merge.
CLIHub (Docker)GitHub · GitLabGo · TypeScript · Python · Rust
Probe Desktop

Watch shared office documents

For finance, legal, sales and operations teams, and anyone who shares files.

  • Follows the Word, Excel and PowerPoint files of your OneDrive or Google Drive folders.
  • Flags a formula replaced by a number, a total that no longer counts every row, an amount or deadline changed, "shall" turned into "may".
  • Keeps a queue of documents to review; once reviewed, a version becomes the reference for the next changes.
WindowsmacOSOneDrive · Google DriveWord · Excel · PowerPoint

Which one do you need?

Probe for codeProbe Desktop
You areA developer, a reviewer, a platform teamA controller, a lawyer, a manager, an assistant
What changesSource code, in Git commits and pull requestsWord, Excel and PowerPoint files, in OneDrive or Google Drive
The referenceThe base branch of the pull requestThe last version someone marked as reviewed
What it looks forSecurity-sensitive code, removed validation, API and dependency changes, failing or missing testsHard-coded formulas, shrunk ranges, changed amounts, dates and obligations, hidden sheets or slides, macros
EvidenceChecks run in a sandbox, reproduced issues, a JSON and Markdown reportThe before and the after of each flagged change, and who saved it last
How it runsA command-line tool in your terminal or CI, or the Hub for a whole accountAn application in the taskbar or the menu bar, watching in the background
Probe Desktop · For documents

Shared files change quietly. Know which changes matter.

A budget passed between five people, a contract reworked by the other party, a board deck updated the night before. Nobody re-reads everything, and nobody should have to. Probe Desktop watches the folders your team already uses and shows you the modifications that change a number, a commitment or a calculation, with the before and the after.

  • Sees what track changes does not. A formula silently replaced by its value, a sum that stops one row short, a sheet made invisible.
  • Works where your files already are. Pick a OneDrive or Google Drive folder synchronized on your computer. No migration, no add-in, no account.
  • Stays on your computer. Documents are compared locally. An AI explanation is optional and receives only the flagged excerpts.
Probe Desktop · 2 documents to review
  • X A range in the formula was reduced: some cells are no longer counted Budget 2026.xlsx · Budget!B1 · high

    =SUM(A1:A3)

    =SUM(A1:A2)

  • X A formula was replaced by its value: the cell no longer updates Budget 2026.xlsx · Budget!D2 · high

    =A2*2

    400

  • W An obligation was softened Supplier contract.docx · Paragraph 2 · high

    The supplier shall deliver before 15/03/2026.

    The supplier may deliver before 15/03/2026.

  • W An amount or a percentage was changed Supplier contract.docx · Paragraph 3 · high

    The price is set at €12,000 excluding VAT.

    The price is set at €15,000 excluding VAT.

Probe for code

AI assistants write large pull requests fast. Probe tells you which lines to read first: it maps the diff to risk signals, runs your checks in disposable containers, tries to reproduce issues with differential tests, and hands you a focused review plan with traceable evidence. A command-line tool for Windows, Linux and macOS, and a Docker hub for a whole GitHub or GitLab account. Try it in 2 minutes →

Probe for code · The process

Know the impact before the code exists. Review only what needs you.

AI-written code gets predictable when the agent announces its work first. Probe measures the risk of the plan before any code is written, then checks that the change did exactly what was announced. A change that follows a low-risk plan and passes its checks can merge without a human reading the diff; everything else goes to review, with the reasons.

  1. Plan and check the intent

    probe plan has the agent simulate the change read-only. Fixed rules, not the model, flag critical paths, public API and dependency changes, widely used or untested code. A flagged plan goes to a human before any code is written.

  2. An agent codes the plan

    The plan becomes a contract: the files, symbols and dependencies the agent may touch. It implements that, and only that.

  3. Check conformance

    probe review --plan re-assesses the plan, runs the checks in the sandbox and compares the diff with the contract: every unplanned file, unannounced API change, critical path or dependency is reported.

  4. Merge, or review

    Conforming change, low-risk plan, checks passed, nothing else flagged: no human review required, exit 0, the pipeline merges. Otherwise exit 2: a human reviews the pull request, starting from the listed reasons.

Exit 0 merge

  1. The plan, re-assessed at review time, raised no risk category and was fully measured.
  2. The diff stays within the plan: no unplanned file, API, critical path or dependency.
  3. Every check passed; no high signal, reproduced issue or unverified area.

Exit 2 human review

  1. The plan touches something critical, public or widely used, or could not be measured.
  2. The change drifted from the plan.
  3. A check failed or the report found something to look at. Every reason is listed.
# 1. before coding: plan, and let a human validate a flagged plan
probe plan --intent-file task.md --ci
# 2. the agent implements .probe/PLAN.json
# 3-4. in CI: conformance, checks and the gate (exit 0 = merge, 2 = review)
probe review --base origin/main --plan .probe/PLAN.json --ci
Plan conformance: conforming; 0 high, 0 medium, 0 low differences from the plan.
Plan gate: no human review required (low-risk plan, conforming change, checks passed, nothing else requests review).

The gate is a process decision, not a proof of correctness: your trusted policy decides what is critical and which tests run, and any doubt sends the change to a human. How the gate decides.

Tests, vet, build and coverage all passed on this change. It also dropped the refund amount validation and let the support role issue refunds. Probe put both at the top of the review plan, among 19 flagged lines out of 125.

Probe for code · Why it saves time

Stop reading AI-generated diffs top to bottom.

The expensive part of reviewing an AI-assisted pull request is not the risky lines. It is finding them among the helpers, tests and docs around them, then working out whether a suspicion is real. Probe does that sorting and gathers the evidence before you open the diff.

Without Probe every line, same attention

  1. Check out the branch and run tests, vet and build by hand.
  2. Read 125 changed lines in file order, helpers first.
  3. Notice, with luck, that a validation call vanished from a refund path.
  4. Write a throwaway test to find out whether it matters.
  5. Approve without a record of what was actually checked.

With Probe risk first, evidence attached

  1. Checks already ran in an isolated container, with logs and hashes.
  2. Open the report: 19 flagged lines, ranked by severity, with reasons.
  3. Start with payment/refund.go, where the report shows the removed validation.
  4. Let the optional reviewer try to reproduce the issue on base and candidate.
  5. Keep the Markdown and JSON reports as the review record.

Less to read first

A deduplicated review surface counted in real changed lines, so large generated changes shrink to the parts that touch authentication, payments, validation, dependencies or public APIs.

No local setup per PR

Your test, typecheck, build and coverage commands run in disposable, network-less containers. Nothing of the candidate executes on your machine.

Know when to stop

Unverified areas and incomplete checks are listed explicitly, and --ci turns them into an exit code. You know what was covered and what still needs your judgment.

Probe for code · Example report

What you open instead of the raw diff

Excerpt of .probe/CONFIDENCE_REPORT.md, produced by probe review HEAD~1..HEAD --reviewer=false --ci (v0.3.0) on a small shop repository. The commit added catalog helpers with tests, reworked refunds and touched authorization. No AI provider was used.

.probe/CONFIDENCE_REPORT.md
## Change Summary
120 additions / 5 deletions · 5 files changed
Exit code: 2. No confidence percentage is assigned.

## Automated Checks1
- PASS test      (check-1; exit 0; 5982 ms)
- PASS typecheck (check-2; exit 0; 4866 ms)
- PASS build     (check-3; exit 0; 2835 ms)
- PASS coverage  (check-4; exit 0; 6416 ms)

## Reproduced Issues2
No issue was reproduced by a passing baseline
and failing candidate experiment.

## Suggested Human Review3
- high auth/auth.go:7 (new): Authentication or
  authorization function body changed
- high payment/refund.go:21–22 (new): Payment-sensitive
  function body changed
- high payment/refund.go:25–26 (new): Exported Go
  declaration added; Payment-sensitive function body changed
- high payment/refund.go:21 (old): Configured sensitive
  path changed; No nearby test file changed;
  Possible input validation removed
- medium catalog/format.go:9 (new): Exported Go
  declaration added
  … 11 more medium signals in catalog/

## Review Surface4
Focused review: 19 / 125 changed lines.

## Changed-line Execution5
Of 67 added Go lines, 44 were executed at least once,
0 were not executed, 23 are not inside any
instrumented block.
  1. Checks are done, and they are not the answer. Every configured command passed in an isolated container. That is exactly why the rest of the report matters.
  2. Evidence, not opinions. A reproduced issue needs a test that passes on the base and fails on the candidate. Without an AI provider nothing is claimed here, and the report says so.
  3. Your reading order. High signals first: an authorization rule and a refund path changed, and a validation call was removed with no nearby test change. That is where the bug is.
  4. The size of the job. 19 of 125 changed lines carry a signal. Read those first, then decide how much of the rest deserves attention.
  5. An observation, not a proof. The new refund code ran during the tests, yet no test asserts the amount check. Probe reports execution but never presents it as proof that the code is tested.
Probe for code

How it works

  1. Compare immutable commits

    Both refs resolve to commit IDs. Renames, deletions, binaries and merge-base semantics are handled; only committed files are reviewed, so local edits never leak into the result.

  2. Collect risk signals

    Go AST comparison, labelled lexical heuristics and sensitive-path rules flag authentication, payments, dependencies, removed validation, unsafe constructs and missing tests in Go, TypeScript/JavaScript, Python and Rust.

  3. Execute in a sandbox

    Your test, typecheck and build commands run as argv arrays in disposable, non-root, network-less containers with CPU, memory, PID and time limits.

  4. Reproduce, don't assert

    An optional reviewer writes temporary tests. A test that passes on the baseline and fails on the candidate supports a reproduced issue; anything else stays UNVERIFIED.

From a first try to every pull request

Try it on your last commit with no configuration, no Docker and no API key. When you adopt it, commit a policy to your base branch: it is read from the base, so a pull request cannot loosen its own rules.

Follow the getting started guide →

# try it: compare with the previous commit
probe lint HEAD~1..HEAD

# adopt it: once, in your repository
probe init
git add .probe.json && git commit -m "Add review policy"

# on every pull request
probe review --base main --ci
Probe for code · Docker

Or run it for a whole account: Probe Hub

The hub is the companion Docker image. Run one container, sign in with GitHub or GitLab, and every repository still missing a .probe.json gets a policy proposed in one click, previewed before it is committed to the default branch.

Repositories that have a policy can be watched: each new commit is analyzed and its report opens on its own. A severity slider filters the alerts from low to critical, and clicking one unfolds every modification it concerns, with the flagged lines highlighted in the diff.

One Docker container, no database, no external service — so a company can deploy it internally against its own GitHub Enterprise or GitLab instance. In its default mode it never executes the analyzed code.

Hosted at app.probe.technology — the same image you can run on your own network.

# one container, on your own network
$ docker run -p 8080:8080 \
    -e PROBE_HUB_BASE_URL=https://hub.example.com \
    -e PROBE_HUB_GITHUB_CLIENT_ID=... \
    -e PROBE_HUB_GITHUB_CLIENT_SECRET=... \
    -v hub-data:/var/lib/probe-hub \
    gvinsot/probe-hub

Sign in · bootstrap the missing policies · watch pushes
Reports open by themselves, filtered by severity.
Probe for code

Capabilities

Focused review surface

Deduplicated review ranges with old and new coordinates, counted in actual changed lines, ranked by severity.

Differential evidence

Generated tests run on baseline and candidate in fresh environments. Go execution is verified from structured test events.

Changed-line execution

Optionally measure which added Go lines a recorded run executed — reported as an observation, never as proof of testing.

Impact analysis

A static index of Go, TypeScript/JavaScript, Python and Rust lists the unchanged callers of each changed function and the existing tests that reach it.

Tests on both revisions

Impacted and changed baseline tests, differential fuzzing and mutation of added lines run on base and candidate, recorded as evidence, never as a verdict.

Plan, then check the scope

probe plan assesses an agent's plan before any code exists; --plan then flags every file or exported API the change touched without announcing it.

Bounded AI reviewer

Bring any OpenAI-compatible provider. Tools are bounded: redacted file reads, search and test runs. No shell, no URL fetcher.

CI ready

A reusable GitHub Actions workflow, meaningful exit codes, JSON reports with commit IDs, evidence and artifact hashes, plus SARIF and a PR comment holding only evidence-backed findings.

Agent friendly

Coding agents run lint and review before handing work over, and report what was reproduced and what remains open.

What Probe does not claim

A review tool earns trust by being precise about its limits. Probe assigns no confidence percentages and never approves a change on a model's word. For code, it lifts human review only through the plan gate, when a low-risk plan was followed and every automated check came back clean. For documents, a person always decides.

  • Signals are reasons to investigate, not confirmed bugs.
  • Passing tests and a small review surface are not correctness guarantees.
  • An executed line is an observation, not proof that it is tested.
  • A model's assertion is never evidence; only a reproduced difference is.
  • The plan gate says the announced, low-risk scope was respected, not that the code is correct.
  • Execution never falls back to your host when Docker is missing.
  • A document marked as reviewed is a human decision, not a guarantee that its figures are right.
  • A document with no finding still changed: the full list of changes stays one click away.

Put your review time where it matters.

Open source, AGPL-3.0 licensed. A command-line tool for code, an application for documents, both published with every release.