OrangePro
UnexploredFind test gaps, generate grounded tests, and dynamically prove behavior with mutation testing.
Install
Terminal
$npx -y @orangepro/mcp-server mcpmcp_config.json
{
"mcpServers": {
"ai-orangepro-mcp": {
"env": {
"OPENAI_API_KEY": "${OPENAI_API_KEY}",
"OLLAMA_BASE_URL": "${OLLAMA_BASE_URL}",
"ANTHROPIC_API_KEY": "${ANTHROPIC_API_KEY}"
},
"args": [
"-y",
"@orangepro/mcp-server",
"mcp"
],
"command": "npx"
}
}
}Documentation
See which code your tests really protect. Prove it by breaking the code on purpose.
OrangePro maps every public function in your repository, links each one to the tests that actually call it, and ranks what's left by what it can break. For a test you care about, it breaks the function in an isolated copy and checks that the test fails. It runs on your machine, uses no AI model for scoring, and sends nothing to OrangePro. Model calls happen only if you add your own key, for optional test generation and suggestions.
npx -y @orangepro/mcp-server@latest start .
Contents
- What you get
- Why developers use it
- Quick start
- Evidence tiers
- How the ranking works
- Prove a test
- Use with your coding agent
- Configuration
- Language support
- Privacy and network use
- Feedback
- Reference
What you get
Every run writes two reports to .orangepro/:
| File | For | What's in it |
|---|---|---|
short_behavior-coverage.html | Leads, reviewers, anyone in a hurry | One page: the headline, where to start, the code paths that delete data and their test evidence, the top-ranked items, the tests behind each proof, and how to reproduce the run. Each name links to its line at the analysed commit (GitHub and GitLab). |
behavior-coverage.html | Developers | The full interactive map: every behavior and its evidence, flows from entry points through services, the ranked list with suggested tests, and the settings used. |
The one-page summary
OrangePro run on its own repository at commit 3cc5c19, with no AI model. The proof at the bottom replaced OrangeProClient.get with a fixed return value in an isolated copy, and the project's own test, unchanged, failed at its assertion.
The detailed report
→ Live example: Twenty CRM (5,237 behaviors mapped)
System map: entry lanes (GraphQL, HTTP, jobs) flowing into services, sized by traffic, colored by evidence tier, red-ringed by risk.
Why developers use it
- It tells you what a test actually checks, not just what it runs. A test that calls a function isn't necessarily checking it.
opro provereplaces the function body with a fixed return value in an isolated copy and reruns your own test, unchanged. If the test still passes, it wasn't protecting that function. - It puts consequences first. Two lists sit above the ranking:
- Can destroy data: code paths that reach a delete or purge with no proven test. Cache evictions don't count.
- Changing fast: code that changes often with nothing proving it works.
- It's honest about evidence. A function is "linked" only when a test calls that exact function. A test with a similar name is a lead, never coverage. Tests that only check mocks don't count.
- It's repeatable. Same commit, full git history, same config and same version give the same ranking. Each report records fingerprints, so you can tell when a change in results came from the code and when it came from the tool.
- Your coding agent can drive it. As an MCP server it gives Claude Code, Cursor, Copilot, Codex and others the ranked gaps, the suggested test location and the proof step in one loop.
- It's free, local and open source (MIT). No account, no API key for analysis, and no upload.
Quick start
cd /path/to/your/repo
# install the repo's own dependencies first (npm ci, uv sync, go mod download, ...)
npx -y @orangepro/mcp-server@latest start .
open .orangepro/short_behavior-coverage.html
Analysis, ranking and proof need no model key. Test generation is optional and uses your own key (see Model setup).
Tips for the most accurate run:
- Use a full clone, not a shallow one. Change history drives part of the ranking, and the report says when history was partial.
- Exclude what isn't product code with
rank_exclude_paths(see Configuration). - Rerun after a change. The detailed report shows what entered, moved up or got resolved since the last run.
What's written:
.orangepro/
├── short_behavior-coverage.html ← one-page summary
├── behavior-coverage.html ← full interactive report
├── graph.json ← the evidence graph (deterministic)
├── ledger.json ← proof certificates
├── rtm.md ← traceability matrix
└── config.json ← optional per-repo settings
orangepro_generated/ ← generated tests (only with a model key); your files are never edited
Evidence tiers
Every behavior gets exactly one tier.
| Tier | What it means |
|---|---|
| Dynamically Proven | A test passed on the original code and failed at its own assertion when this function was broken in an isolated copy. |
| Runtime-covered | A coverage tool you ran executed this code. |
| Statically Linked | A test calls this exact function. That's a structural link, not proof. |
| Name match only | A test with a similar name exists. That's a lead, not evidence. |
| No test found | Nothing links a test to it. |
What OrangePro doesn't count, so linked numbers are a floor, not a coverage percentage:
- tests that only assert on mocks;
- calls made over HTTP or from end-to-end suites;
- in Python, objects that reach a test only through a fixture.
Check the tests before writing new ones for a flagged path.
"Dynamically Proven 0" is normal on a first run. Proof runs your tests, so it happens only for the functions you choose, or within the attempt budget of
opro start.
How the ranking works
Each unproven function gets an OrangePro Risk Score, ORS = P × I × D:
- P: how likely it is to change. Change history, fan-out, new code and size.
- I: what it can break. Incoming references, entry-point position, data sensitivity, and a floor for paths that reach a delete or purge within two calls.
- D: how hard a break would be to notice. Evidence tier, and whether it runs unattended (jobs and schedulers).
The score sets the order of work. It is not a defect probability. The weights are fixed and no AI model is involved. The report shows the inputs behind every row.
Prove a test
# Python: the mutation value is derived from the function (return annotation or its observed result)
opro prove-loop --target-symbol 'sym:app/billing/invoices.py#void_invoice' \
--test 'tests/test_invoices.py::test_void_marks_invoice_void' --replacement sentinel
# TypeScript / JavaScript: give the inert body to substitute
opro prove-loop --target-symbol 'sym:src/orders.service.ts#OrdersService.cancel' \
--test src/orders.service.spec.ts --replacement 'return null;'
- Proven: the test passes on the original code and fails at its own assertion on the broken copy.
- Not proven: the test still passes on the broken copy, so it doesn't protect that function. That's a finding too.
- Unrunnable: setup failed. This is never counted either way.
Find weak tests without a model key:
opro roast . # passing tests whose targeted mutant still survives
Use with your coding agent
OrangePro runs as an MCP server. Add it to your client's config:
{
"mcpServers": {
"orangepro-local": {
"command": "npx",
"args": ["-y", "@orangepro/mcp-server@latest", "mcp"]
}
}
}
| Client | Where to put it |
|---|---|
| Claude Code | .mcp.json or ~/.claude.json |
| Cursor | ~/.cursor/mcp.json or Settings → MCP |
| VS Code / Copilot | MCP settings |
| Codex / OpenCode / Windsurf | npx -y @orangepro/mcp-server@latest agent --client codex prints the setup |
A prompt that runs the whole loop:
"Use
orangepro_start, thenorangepro_find_test_gaps. For the top gap, write a test at the suggested path, run it, then callorangepro_prove_loopand tell me whether it was proven."
Configuration
Optional. Put it in .orangepro/config.json in the repository, or in ~/.orangepro/config.json for defaults across repositories. Every setting that changes the ranking is shown in the report.
{
"classification": {
"rank_exclude_paths": ["ui/**", "docs/**", "scripts/**"],
"test_support_paths": ["internal/testutil/**"],
"destructive_sinks": ["archive*"]
},
"tuning": { "churn_window_days": 180 },
"proof": { "python_runner": "auto", "attempt_limit": 20 },
"overrides": [
{ "symbol": "sym:src/legacy/shim.ts#shim", "action": "suppress", "reason": "generated shim, not product code" }
]
}
rank_exclude_pathsremoves paths from the ranking. They are still counted as behaviors.destructive_sinksadds delete-like calls for your codebase.- Overrides (
suppress,pin,reclassify) each require a reason, and the report lists them. - Weights and tiers can't be configured.
Language support
| Language | Map and link | Generated tests | Mutation proof |
|---|---|---|---|
| TypeScript / JavaScript | ✓ | ✓ Jest / Vitest / Mocha | ✓ |
| Python | ✓ | ✓ pytest | ✓ pytest |
| Go | ✓ | ✓ *_test.go | ✓ |
| Java | ✓ | ✓ JUnit 4/5 | ✓ |
| Kotlin, Rust, PHP, C#, Ruby, Swift, C, C++ | ✓ | planned | planned |
Mapping uses tree-sitter. Proof is deliberately narrower: each language needs a runner, a mutation locator and a sandbox profile.
Privacy and network use
- Nothing about your code or your runs is sent anywhere. There is no usage telemetry.
- Source is read in-process and never stored or uploaded. Reports and the evidence graph contain metadata, not your source.
- Your source files are never edited. Proofs run in an isolated copy.
- Model calls happen only if you configure a key, and they go directly from your machine to the provider you chose. Keys are read from the environment and never written to disk.
- Reports load nothing from the network. The only outbound links are ones you click.
Feedback
Reports include a Give feedback link and a This looks wrong link on each finding. Once in a while, after the findings, they also ask whether the report helped.
- The links carry nothing about your project. The form sends only what you review and submit.
- You stay anonymous unless you leave an email.
opro feedback # print the link and your settings
opro feedback off # stop the "Did this help?" question (links stay)
ORANGEPRO_FEEDBACK_URL=off # hide every feedback link
The question appears at most once every 14 days (ORANGEPRO_FEEDBACK_COOLDOWN_DAYS). That preference is stored only on your machine. MCP clients get the link as optional result metadata (_meta["ai.orangepro/feedback_url"]), and your agent is never asked to prompt you for feedback.
Reference
CLI
opro # same as opro start .
opro start . --no-ai --no-auto # analyze + reports, no model calls, no proof attempts
opro start --base main # scope to a branch diff
opro analyze # build the evidence graph and reports
opro gaps --limit 10 # top unproven behaviors, ranked
opro prove-loop ... # mutation proof for one function (see above)
opro roast . # find tests whose mutant survives (no key needed)
opro doctor --proof # why top targets aren't proven yet
opro generate --base main # tests for what this branch changed (needs a model key)
opro rtm # traceability matrix
opro export # metadata-only evidence pack
opro coverage # find or generate runtime coverage artifacts
opro feedback [on|off] # feedback link and invitation setting
opro mcp # run as an MCP server (stdio)
Add --json to any read command for machine output. Run opro help for every flag.
MCP tools (18)
| Tool | What it does |
|---|---|
orangepro_start | Analyze, write both reports, return next actions |
orangepro_analyze_sources | Build or refresh the evidence graph |
orangepro_find_test_gaps | Unproven behaviors ranked by ORS |
orangepro_prove_loop | Setup + mutation proof + report refresh for one behavior |
orangepro_prove | Mutation proof only |
orangepro_generate_tests | Grounded tests for gaps (model key required) |
orangepro_changed_impact | What a diff touches |
orangepro_status | Workspace state without running anything |
orangepro_doctor | What evidence to add next |
orangepro_graph_score | Graph readiness (0–100) |
orangepro_rtm | Traceability matrix |
orangepro_stats | Aggregate statistics |
orangepro_record_run | Record a test run result |
orangepro_explain_test | Why a test was generated |
orangepro_export_evidence_pack | Metadata-only evidence pack |
orangepro_update_graph | Incremental graph update |
orangepro_ai_links | Weak behavior→code suggestions (optional AI, never evidence) |
orangepro_ai_flows | Candidate flows (optional AI, never evidence) |
Coverage artifacts
Run your own unit and integration coverage first, then opro start. OrangePro ingests Go coverprofiles, lcov, coverage.py XML and JaCoCo XML, and keeps unit, integration and unclassified coverage separate. To label artifacts, add .orangepro/coverage-suites.json:
{
"artifacts": {
".orangepro/coverage/unit.coverprofile": { "suite": "unit", "command": "make unit-test-coverage" },
".orangepro/coverage/integration.coverprofile": { "suite": "integration", "command": "make integration-test-coverage" }
}
}
Model setup (BYOK, only for generation)
| Provider | Environment variable |
|---|---|
| OpenAI-compatible | OPENAI_API_KEY (optional: OPENAI_BASE_URL, OPENAI_MODEL) |
| Anthropic | ANTHROPIC_API_KEY (optional: ANTHROPIC_MODEL) |
| Ollama (local, no key) | OLLAMA_BASE_URL (optional: OLLAMA_MODEL) |
Auto-detect order: OpenAI → Ollama → Anthropic. Override with --provider and --model, or run opro setup.
- Model output never changes an evidence tier. Only the mutation proof can mark a behavior Proven.
- Generated tests land in
orangepro_generated/with the files and existing tests they're grounded on, and where and how to run them. - Generated tests that can't run yet are kept as manual steps, with the blocker named.
PR workflow
opro generate --base main # tests for what this branch changed (read-only git diff)
opro generate --changed # current branch vs its base
opro generate --pr 1234 # checks out PR #1234 (asks first; refuses on a dirty tree)
Hosted platform
This repository is the free local tool. The OrangePro platform adds:
- a persistent graph across PRs and repositories;
- CI gates on evidence and risk changes;
- requirement and incident correlation;
- team dashboards.
Contributing
git clone https://github.com/OrangeproAI/orangepro-mcp.git
cd orangepro-mcp && npm ci && npm run build
npm test
PRs welcome. Please open an issue first for large changes.
MIT License · orangepro.ai
Sourced from the repository README.
More in AI & Agents
- PonytailMakes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.109,599
- AgentsMulti-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, and Google Antigravity39,079
- Frontend SlidesCreate beautiful slides on the web using a coding agent's frontend skills28,060
- Agent Skills Search ServerSearch and discover Agent Skills from the skills.sh registry. Powered by HAPI MCP server.25,980
- Agency Agents Zh🎭 267 个即插即用的 AI 专家角色 — 支持 Hermes Agent/Claude Code/Cursor/Copilot 等 18 种工具,覆盖工程/设计/营销/金融等 20 个部门。含 52 个中国市场原创智能体(小红书/抖音/微信/飞书/钉钉等)。搭配编排器 agency-orchestrator,一句话即可让多位专家按 DAG 自动协作。19,868
- Watermarks RemoverStrip multi-vendor AI provenance marks: Unicode text hygiene, statistical rewrite hooks, and C2PA/metadata from PNG/JPEG/SVG/PDF/DOCX/HTML/MD17,822