Search DevTools

Jump to any tool or page

OrangePro

Unexplored

Find test gaps, generate grounded tests, and dynamically prove behavior with mutation testing.

OrangeproAI18 stars1 forksAI & Agents
View source

Install

Terminal

$npx -y @orangepro/mcp-server mcp

mcp_config.json

{
  "mcpServers": {
    "ai-orangepro-mcp": {
      "env": {
        "OPENAI_API_KEY": "${OPENAI_API_KEY}",
        "OLLAMA_BASE_URL": "${OLLAMA_BASE_URL}",
        "ANTHROPIC_API_KEY": "${ANTHROPIC_API_KEY}"
      },
      "args": [
        "-y",
        "@orangepro/mcp-server",
        "mcp"
      ],
      "command": "npx"
    }
  }
}

Documentation

See which code your tests really protect. Prove it by breaking the code on purpose.


OrangePro maps every public function in your repository, links each one to the tests that actually call it, and ranks what's left by what it can break. For a test you care about, it breaks the function in an isolated copy and checks that the test fails. It runs on your machine, uses no AI model for scoring, and sends nothing to OrangePro. Model calls happen only if you add your own key, for optional test generation and suggestions.

npx -y @orangepro/mcp-server@latest start .

Contents


What you get

Every run writes two reports to .orangepro/:

FileForWhat's in it
short_behavior-coverage.htmlLeads, reviewers, anyone in a hurryOne page: the headline, where to start, the code paths that delete data and their test evidence, the top-ranked items, the tests behind each proof, and how to reproduce the run. Each name links to its line at the analysed commit (GitHub and GitLab).
behavior-coverage.htmlDevelopersThe full interactive map: every behavior and its evidence, flows from entry points through services, the ranked list with suggested tests, and the settings used.

The one-page summary

OrangePro run on its own repository at commit 3cc5c19, with no AI model. The proof at the bottom replaced OrangeProClient.get with a fixed return value in an isolated copy, and the project's own test, unchanged, failed at its assertion.

The detailed report

→ Live example: Twenty CRM (5,237 behaviors mapped)

System map: entry lanes (GraphQL, HTTP, jobs) flowing into services, sized by traffic, colored by evidence tier, red-ringed by risk.


Why developers use it

  • It tells you what a test actually checks, not just what it runs. A test that calls a function isn't necessarily checking it. opro prove replaces the function body with a fixed return value in an isolated copy and reruns your own test, unchanged. If the test still passes, it wasn't protecting that function.
  • It puts consequences first. Two lists sit above the ranking:
    • Can destroy data: code paths that reach a delete or purge with no proven test. Cache evictions don't count.
    • Changing fast: code that changes often with nothing proving it works.
  • It's honest about evidence. A function is "linked" only when a test calls that exact function. A test with a similar name is a lead, never coverage. Tests that only check mocks don't count.
  • It's repeatable. Same commit, full git history, same config and same version give the same ranking. Each report records fingerprints, so you can tell when a change in results came from the code and when it came from the tool.
  • Your coding agent can drive it. As an MCP server it gives Claude Code, Cursor, Copilot, Codex and others the ranked gaps, the suggested test location and the proof step in one loop.
  • It's free, local and open source (MIT). No account, no API key for analysis, and no upload.

Quick start

cd /path/to/your/repo
# install the repo's own dependencies first (npm ci, uv sync, go mod download, ...)
npx -y @orangepro/mcp-server@latest start .
open .orangepro/short_behavior-coverage.html

Analysis, ranking and proof need no model key. Test generation is optional and uses your own key (see Model setup).

Tips for the most accurate run:

  • Use a full clone, not a shallow one. Change history drives part of the ranking, and the report says when history was partial.
  • Exclude what isn't product code with rank_exclude_paths (see Configuration).
  • Rerun after a change. The detailed report shows what entered, moved up or got resolved since the last run.

What's written:

.orangepro/
├── short_behavior-coverage.html   ← one-page summary
├── behavior-coverage.html         ← full interactive report
├── graph.json                     ← the evidence graph (deterministic)
├── ledger.json                    ← proof certificates
├── rtm.md                         ← traceability matrix
└── config.json                    ← optional per-repo settings

orangepro_generated/               ← generated tests (only with a model key); your files are never edited

Evidence tiers

Every behavior gets exactly one tier.

TierWhat it means
Dynamically ProvenA test passed on the original code and failed at its own assertion when this function was broken in an isolated copy.
Runtime-coveredA coverage tool you ran executed this code.
Statically LinkedA test calls this exact function. That's a structural link, not proof.
Name match onlyA test with a similar name exists. That's a lead, not evidence.
No test foundNothing links a test to it.

What OrangePro doesn't count, so linked numbers are a floor, not a coverage percentage:

  • tests that only assert on mocks;
  • calls made over HTTP or from end-to-end suites;
  • in Python, objects that reach a test only through a fixture.

Check the tests before writing new ones for a flagged path.

"Dynamically Proven 0" is normal on a first run. Proof runs your tests, so it happens only for the functions you choose, or within the attempt budget of opro start.


How the ranking works

Each unproven function gets an OrangePro Risk Score, ORS = P × I × D:

  • P: how likely it is to change. Change history, fan-out, new code and size.
  • I: what it can break. Incoming references, entry-point position, data sensitivity, and a floor for paths that reach a delete or purge within two calls.
  • D: how hard a break would be to notice. Evidence tier, and whether it runs unattended (jobs and schedulers).

The score sets the order of work. It is not a defect probability. The weights are fixed and no AI model is involved. The report shows the inputs behind every row.


Prove a test

# Python: the mutation value is derived from the function (return annotation or its observed result)
opro prove-loop --target-symbol 'sym:app/billing/invoices.py#void_invoice' \
  --test 'tests/test_invoices.py::test_void_marks_invoice_void' --replacement sentinel

# TypeScript / JavaScript: give the inert body to substitute
opro prove-loop --target-symbol 'sym:src/orders.service.ts#OrdersService.cancel' \
  --test src/orders.service.spec.ts --replacement 'return null;'
  • Proven: the test passes on the original code and fails at its own assertion on the broken copy.
  • Not proven: the test still passes on the broken copy, so it doesn't protect that function. That's a finding too.
  • Unrunnable: setup failed. This is never counted either way.

Find weak tests without a model key:

opro roast .   # passing tests whose targeted mutant still survives

Use with your coding agent

OrangePro runs as an MCP server. Add it to your client's config:

{
  "mcpServers": {
    "orangepro-local": {
      "command": "npx",
      "args": ["-y", "@orangepro/mcp-server@latest", "mcp"]
    }
  }
}
ClientWhere to put it
Claude Code.mcp.json or ~/.claude.json
Cursor~/.cursor/mcp.json or Settings → MCP
VS Code / CopilotMCP settings
Codex / OpenCode / Windsurfnpx -y @orangepro/mcp-server@latest agent --client codex prints the setup

A prompt that runs the whole loop:

"Use orangepro_start, then orangepro_find_test_gaps. For the top gap, write a test at the suggested path, run it, then call orangepro_prove_loop and tell me whether it was proven."


Configuration

Optional. Put it in .orangepro/config.json in the repository, or in ~/.orangepro/config.json for defaults across repositories. Every setting that changes the ranking is shown in the report.

{
  "classification": {
    "rank_exclude_paths": ["ui/**", "docs/**", "scripts/**"],
    "test_support_paths": ["internal/testutil/**"],
    "destructive_sinks": ["archive*"]
  },
  "tuning": { "churn_window_days": 180 },
  "proof": { "python_runner": "auto", "attempt_limit": 20 },
  "overrides": [
    { "symbol": "sym:src/legacy/shim.ts#shim", "action": "suppress", "reason": "generated shim, not product code" }
  ]
}
  • rank_exclude_paths removes paths from the ranking. They are still counted as behaviors.
  • destructive_sinks adds delete-like calls for your codebase.
  • Overrides (suppress, pin, reclassify) each require a reason, and the report lists them.
  • Weights and tiers can't be configured.

Language support

LanguageMap and linkGenerated testsMutation proof
TypeScript / JavaScript✓✓ Jest / Vitest / Mocha✓
Python✓✓ pytest✓ pytest
Go✓✓ *_test.go✓
Java✓✓ JUnit 4/5✓
Kotlin, Rust, PHP, C#, Ruby, Swift, C, C++✓plannedplanned

Mapping uses tree-sitter. Proof is deliberately narrower: each language needs a runner, a mutation locator and a sandbox profile.


Privacy and network use

  • Nothing about your code or your runs is sent anywhere. There is no usage telemetry.
  • Source is read in-process and never stored or uploaded. Reports and the evidence graph contain metadata, not your source.
  • Your source files are never edited. Proofs run in an isolated copy.
  • Model calls happen only if you configure a key, and they go directly from your machine to the provider you chose. Keys are read from the environment and never written to disk.
  • Reports load nothing from the network. The only outbound links are ones you click.

Feedback

Reports include a Give feedback link and a This looks wrong link on each finding. Once in a while, after the findings, they also ask whether the report helped.

  • The links carry nothing about your project. The form sends only what you review and submit.
  • You stay anonymous unless you leave an email.
opro feedback          # print the link and your settings
opro feedback off      # stop the "Did this help?" question (links stay)
ORANGEPRO_FEEDBACK_URL=off   # hide every feedback link

The question appears at most once every 14 days (ORANGEPRO_FEEDBACK_COOLDOWN_DAYS). That preference is stored only on your machine. MCP clients get the link as optional result metadata (_meta["ai.orangepro/feedback_url"]), and your agent is never asked to prompt you for feedback.


Reference

CLI

opro                          # same as opro start .
opro start . --no-ai --no-auto   # analyze + reports, no model calls, no proof attempts
opro start --base main        # scope to a branch diff
opro analyze                  # build the evidence graph and reports
opro gaps --limit 10          # top unproven behaviors, ranked
opro prove-loop ...           # mutation proof for one function (see above)
opro roast .                  # find tests whose mutant survives (no key needed)
opro doctor --proof           # why top targets aren't proven yet
opro generate --base main     # tests for what this branch changed (needs a model key)
opro rtm                      # traceability matrix
opro export                   # metadata-only evidence pack
opro coverage                 # find or generate runtime coverage artifacts
opro feedback [on|off]        # feedback link and invitation setting
opro mcp                      # run as an MCP server (stdio)

Add --json to any read command for machine output. Run opro help for every flag.

MCP tools (18)

ToolWhat it does
orangepro_startAnalyze, write both reports, return next actions
orangepro_analyze_sourcesBuild or refresh the evidence graph
orangepro_find_test_gapsUnproven behaviors ranked by ORS
orangepro_prove_loopSetup + mutation proof + report refresh for one behavior
orangepro_proveMutation proof only
orangepro_generate_testsGrounded tests for gaps (model key required)
orangepro_changed_impactWhat a diff touches
orangepro_statusWorkspace state without running anything
orangepro_doctorWhat evidence to add next
orangepro_graph_scoreGraph readiness (0–100)
orangepro_rtmTraceability matrix
orangepro_statsAggregate statistics
orangepro_record_runRecord a test run result
orangepro_explain_testWhy a test was generated
orangepro_export_evidence_packMetadata-only evidence pack
orangepro_update_graphIncremental graph update
orangepro_ai_linksWeak behavior→code suggestions (optional AI, never evidence)
orangepro_ai_flowsCandidate flows (optional AI, never evidence)

Coverage artifacts

Run your own unit and integration coverage first, then opro start. OrangePro ingests Go coverprofiles, lcov, coverage.py XML and JaCoCo XML, and keeps unit, integration and unclassified coverage separate. To label artifacts, add .orangepro/coverage-suites.json:

{
  "artifacts": {
    ".orangepro/coverage/unit.coverprofile": { "suite": "unit", "command": "make unit-test-coverage" },
    ".orangepro/coverage/integration.coverprofile": { "suite": "integration", "command": "make integration-test-coverage" }
  }
}

Model setup (BYOK, only for generation)

ProviderEnvironment variable
OpenAI-compatibleOPENAI_API_KEY (optional: OPENAI_BASE_URL, OPENAI_MODEL)
AnthropicANTHROPIC_API_KEY (optional: ANTHROPIC_MODEL)
Ollama (local, no key)OLLAMA_BASE_URL (optional: OLLAMA_MODEL)

Auto-detect order: OpenAI → Ollama → Anthropic. Override with --provider and --model, or run opro setup.

  • Model output never changes an evidence tier. Only the mutation proof can mark a behavior Proven.
  • Generated tests land in orangepro_generated/ with the files and existing tests they're grounded on, and where and how to run them.
  • Generated tests that can't run yet are kept as manual steps, with the blocker named.

PR workflow

opro generate --base main     # tests for what this branch changed (read-only git diff)
opro generate --changed       # current branch vs its base
opro generate --pr 1234       # checks out PR #1234 (asks first; refuses on a dirty tree)

Hosted platform

This repository is the free local tool. The OrangePro platform adds:

  • a persistent graph across PRs and repositories;
  • CI gates on evidence and risk changes;
  • requirement and incident correlation;
  • team dashboards.

Contributing

git clone https://github.com/OrangeproAI/orangepro-mcp.git
cd orangepro-mcp && npm ci && npm run build
npm test

PRs welcome. Please open an issue first for large changes.


MIT License · orangepro.ai

Sourced from the repository README.

More in AI & Agents