Features
Nine features, one question: what actually happened in that run?
Last updated:
The capabilities
Each one answers a question a review already asks
A contract that says what done means
Describe the change in one sentence. Reality Graph proposes the goal, the paths and the checks; you correct and confirm them before anything runs.
- Empty acceptance criteria fail validation
- Eleven declared intents, never guessed
- A pre-run recommendation with no authority to block
Context bounded by that contract
Your coding tool receives the files the contract declares, in a fixed order, with credentials removed before the prompt is assembled.
- No retrieval and no similarity search
- .env refused in every case
- Every token count labelled exact or estimated
A local index you can check
One node per module, class, function and method, anchored to file and line, with a content hash that tells you whether the map still matches.
- Python only, other languages index to nothing
- Five standard-library imports, no network
- Fresh, stale or incomplete, by content hash
Checks Reality Graph runs itself
Approved commands run as an argv list with no shell and a bounded timeout, and the observed exit code is recorded rather than reported.
- Executed or attested, with no third value
- A missing result blocks instead of passing
- No sandbox, and the page says so
One verdict, computed the same way every time
Verified, verified with limitations or blocked, from one deterministic function fed by changed files read out of Git.
- Green tests can still end blocked
- One call site, so CLI and dashboard agree
- Receipts carry a digest that flags later edits
Every run, still open
One local workspace on 127.0.0.1 holding every run, review, check and receipt, rendering exactly what the CLI prints.
- Eight views, one canonical projection
- The browser never computes a verdict
- A port per worktree, and no request off the machine
Every prompt becomes a ticket
Your backlog and your agent's work are the same record: a command runs a ticket, and a prompt without one gets one.
- The mission lives inside the ticket
- No status moves without evidence
- Append-only, so failed attempts stay legible
Four tools your agent gets, ten it does not
A local read-only MCP server over stdio. The dangerous operations are not missing from the contract, they are in it and refused by name.
- No port, no socket, no network call
- approve_own_patch is declared and denied
- Deterministic ranking, no embeddings
Approvals that do not follow the branch
A protected path that changed forces a block nothing clears, and an approval is bound to the run, HEAD, scope and checks it was given for.
- Any of those four moving marks it stale
- A local attestation, not an identity
- Detection after the fact, not prevention
By design
The list an evaluator reads first
You can rely on
- A contract validated before the run, with acceptance criteria that cannot be left empty.
- Context assembled only from what that contract declares, with credentials scrubbed first.
- Checks executed by Reality Graph, with the exit code it observed.
- Changed files and diffs read from Git, not from the model's prose.
- One deterministic verdict, and a local record that outlives the session.
Not true today
- Team features. No accounts, no seats, no roles, no shared history.
- CI or pull-request integration. No Git hooks, no GitHub App, no pipeline import.
- Test-strength analysis. No assertion counting, no coverage, no mutation testing.
- Indexing anything but Python, and no incremental rebuild.
- Any certification, and no benchmark or token-saving figure.
Validated on Linux with Python 3.12 and the official Codex CLI. Windows and macOS are unverified in the current release matrix.
How the six pieces fit into one run is walked through on the product page, and what leaves your machine is traced on the data-path page.
Questions people actually ask
- Is the verdict an opinion from a model?
- No. It is computed from the evidence a run collected, by rules that give the same answer twice on the same inputs. That is why there are three outcomes and no score: a number invites you to read confidence into it, and there is no model judgement here to be confident about.
- Does it work on a codebase that is not Python?
- Partly. The mission contract, the executed checks, the evidence from Git, the verdict and the approvals are language-agnostic. CodeGraph is not: it reads Python only, and any other repository produces an empty graph.
- Can our whole team use this?
- Not as a team product. Reality Graph is a single-operator, single-machine tool with no accounts, no seats and no shared state, so there are no roles, no team approvals and no history a colleague can open. Several people can run it, but each of them runs their own.
- Do I have to use all of this to get anything out of it?
- No, but these are stages of one run rather than modules you switch on. A mission that demands little produces a thin verdict; one that demands more produces a verdict with more behind it. What you vary is how much a run has to establish, not which parts of the product exist.
- What is the smallest useful way to start?
- One project and one task. The loop does not need a rollout: your task becomes a checked mission, the approved checks run, and you get a verdict with the evidence behind it. Nothing depends on your team adopting it first, because there are no accounts and no shared state to set up. What it will not do is judge a code base it has never run on. The record starts when you start.
Pick the feature closest to the argument you keep having
Every page states what its feature does, and where it stops, in the same breath.