OrangePro
UnexploredFind test gaps, generate grounded tests, and dynamically prove behavior with mutation testing.
Install
Terminal
$npx -y @orangepro/mcp-server mcpmcp_config.json
{
"mcpServers": {
"ai-orangepro-mcp": {
"env": {
"OPENAI_API_KEY": "${OPENAI_API_KEY}",
"OLLAMA_BASE_URL": "${OLLAMA_BASE_URL}",
"ANTHROPIC_API_KEY": "${ANTHROPIC_API_KEY}"
},
"args": [
"-y",
"@orangepro/mcp-server",
"mcp"
],
"command": "npx"
}
}
}Documentation
Find the behaviors your tests miss. Generate grounded tests that actually run.
OrangePro maps every public behavior in your codebase, scores each one by real test evidence, and shows you the structural blind spots before your users find them. Runs locally. Your code never leaves your machine.
npx -y @orangepro/mcp-server@latest start .
Table of Contents
- What you get
- Evidence tiers
- Quick start
- Use with your coding agent
- How it works
- Language support
- Privacy
- CLI reference
- MCP tools
- Platform
- Contributing
What you get
One command produces an interactive HTML report:
npx -y @orangepro/mcp-server@latest start .
open .orangepro/behavior-coverage.html
The report has two modes: Simple (integration-level blind spots, plain English) and Expert (full behavior list, evidence tiers, flows, system map). Toggle with the pill switch at the top.
โ Live example: Twenty CRM (5,237 behaviors mapped)
System map โ entry lanes (GraphQL, HTTP, Jobs) flowing into services, sized by traffic, colored by evidence tier, red-ringed by risk.
Priority gaps of another open source Project HONO โ top 20 unproven behaviors ranked by blast radius, with generated test drafts.
Evidence tiers
Every behavior gets exactly one tier. Nothing is labeled "tested" on faith.
| Tier | Color | What it means |
|---|---|---|
| Dynamically Proven | ๐ข | A real test kills a targeted mutation of this behavior |
| Runtime-covered | ๐ข | Coverage tool executed this code |
| Statically Linked | ๐ก | A test imports and calls this code โ structural link, not proof |
| Unconfirmed Candidate | โช | A similar test file exists โ a lead, not evidence |
| No Signal | ๐ด | Nothing tests this behavior |
"Dynamically Proven 0" is normal on first run. Proof requires running tests against targeted mutations. That's the trust model.
Quick start
cd /path/to/your/repo
npm install # install the repo's own dependencies first
npx -y @orangepro/mcp-server@latest start .
open .orangepro/behavior-coverage.html
No API key needed. The report shows your system map, evidence tiers, priority gaps, and delta since last run.
Want test generation? Add a model key (BYOK):
export ANTHROPIC_API_KEY="..." # or OPENAI_API_KEY / OLLAMA_BASE_URL
npx -y @orangepro/mcp-server@latest start .
AI output never changes evidence tiers. Only the mutation-kill oracle can mint Dynamically Proven.
Output:
.orangepro/
โโโ behavior-coverage.html โ open this
โโโ graph.json โ deterministic evidence graph
โโโ COVERAGE_REPORT.md โ coverage and gap summary
โโโ ai/ โ candidate flows (when a key is configured)
orangepro_generated/ โ generated tests; your source files are never touched
Each rerun shows a delta banner: what entered the codebase, what moved up in risk, what got resolved.
Use with your coding agent
OrangePro runs as an MCP server. Add to your client's config:
{
"mcpServers": {
"orangepro-local": {
"command": "npx",
"args": ["-y", "@orangepro/mcp-server@latest", "mcp"]
}
}
}
| Client | Where to put it |
|---|---|
| Claude Code | .mcp.json or ~/.claude.json |
| Cursor | ~/.cursor/mcp.json or Settings โ MCP |
| VS Code / Copilot | MCP settings |
| Codex / OpenCode | Run npx -y @orangepro/mcp-server@latest agent --client codex |
The workflow: Tell your agent:
"Use
orangepro_start, thenorangepro_generate_testswith base_ref=main. Write each test to its suggested_path, run it, and report pass/fail."
The agent writes the test, runs it, calls orangepro_prove, and the behavior turns Dynamically Proven. One prompt, full loop.
Works with
Claude Code ยท Cursor ยท GitHub Copilot ยท Codex ยท Windsurf ยท OpenCode ยท VS Code
Any MCP-compatible agent can drive OrangePro. No vendor lock-in.
How it works
โโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโ
โ Your Code โ โโโบ โ Knowledge โ โโโบ โ Evidence โ
โ (any lang) โ โ Graph โ โ Tiers โ
โโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโ
โ
โโโโโโโโดโโโโโโโ
โผ โผ
โโโโโโโโโโโโโ โโโโโโโโโโโโ
โ Gap Reportโ โ Generate โ
โ + Risks โ โ Tests โ
โโโโโโโโโโโโโ โโโโโโโโโโโโ
| Phase | What happens | Needs a model key? |
|---|---|---|
| Analyze | AST walk โ behaviors, flows, evidence tiers | No |
| Score | Graph readiness score (0โ100) | No |
| Generate | Grounded tests for top gaps | Yes (BYOK) |
| Prove | Mutation-kill oracle confirms test breaks if behavior changes | No |
Same code = same score. Deterministic. Always.
Language support
| Language | Static mapping | Generated tests | Dynamic proof |
|---|---|---|---|
| TypeScript / JavaScript | โ | โ Jest / Vitest / Mocha | โ |
| Python | โ | โ pytest | โ |
| Go | โ | โ *_test.go | โ |
| Java | โ | โ JUnit 4/5 | โ |
| Kotlin, Rust, PHP, C#, Ruby, Swift, C, C++ | โ | planned | planned |
Static mapping works across many languages via tree-sitter. Dynamic proof is deliberately narrower โ each language needs a runner, mutation locator, and sandbox profile.
Privacy
- No stored source. Reads code in-process. Never uploads to an OrangePro server.
- No existing-source mutation. Never edits your source or test files.
- Your keys stay yours. Read from env at call time, never persisted.
- BYOK is direct. Code context goes to the model provider you configure. OrangePro is not in that path.
CLI reference
opro # analyze + report + agent next actions
opro start --base main # same, scoped to a branch diff
opro analyze # build the evidence graph
opro score # graph readiness (0โ100)
opro gaps --limit 10 # top 10 untested behaviors
opro generate --base main # tests for PR diff
opro generate --single # top gap, whole repo
opro prove # mutation-kill oracle
opro rtm # traceability matrix
opro export # metadata-only evidence pack
opro mcp # run as MCP server (stdio)
opro doctor # what evidence to add next
opro coverage # ingest runtime coverage
Add --json to any read command for machine output. Run opro help for the full reference.
MCP tools (18 total)
| Tool | What it does |
|---|---|
orangepro_start | One-command setup: analyze + report + next actions |
orangepro_analyze_sources | Build/refresh the evidence graph |
orangepro_generate_tests | Generate grounded tests for gaps |
orangepro_prove | Run mutation-kill oracle on a behavior |
orangepro_prove_loop | Setup + dynamic proof + report refresh for one behavior |
orangepro_find_test_gaps | List behaviors with weak/missing tests, ranked by risk |
orangepro_graph_score | Graph readiness score (0โ100) |
orangepro_status | Workspace state without generating anything |
orangepro_doctor | Recommend next evidence to improve quality |
orangepro_rtm | Requirements traceability matrix |
orangepro_stats | Aggregate statistics |
orangepro_changed_impact | What a diff touches (requires git + base ref) |
orangepro_record_run | Record a test run result |
orangepro_explain_test | Explain why a test was generated |
orangepro_export_evidence_pack | Export metadata-only evidence pack |
orangepro_update_graph | Incremental graph update |
orangepro_ai_links | Weak behaviorโsymbol suggestions (optional AI) |
orangepro_ai_flows | Candidate flow discovery (optional AI) |
PR workflow
opro generate --base main # tests for what this branch changed
opro generate --pr 1234 # checks out PR #1234
opro generate --changed # current branch diff vs main
Each generated test includes:
- Grounding โ the real files, symbols, and existing tests it cites
- Run hints โ where to write it, how to run it
- Scenario bucket โ what failure mode it targets
If dependencies aren't installed, tests are kept as Manual tests (Given/When/Then steps with the blocker named). Install dependencies and re-run to convert them to runnable tests.
Test categories
Generation is evidence-gated. A category is produced only when the graph has supporting evidence.
| Category | What it targets |
|---|---|
| Happy path | Primary expected behavior |
| Validation error | Bad/invalid input handling |
| Edge case | Boundaries, empty/null, concurrency, retries |
| Integration flow | Multi-step behavior across services |
| Security / privacy | Auth, injection, data leakage |
| Regression | Pinning a previously-broken behavior |
Model setup (BYOK)
Analysis, scoring, and proof need no model key. Generation does.
| Provider | Environment variable |
|---|---|
| OpenAI-compatible | OPENAI_API_KEY (optional: OPENAI_BASE_URL, OPENAI_MODEL) |
| Anthropic | ANTHROPIC_API_KEY (optional: ANTHROPIC_MODEL) |
| Ollama (local, no key) | OLLAMA_BASE_URL (optional: OLLAMA_MODEL) |
Auto-detect order: OpenAI โ Ollama โ Anthropic. Override with --provider and --model.
Run opro setup to configure interactively. Keys stay in your environment โ never written to graph, config, or artifacts.
AI candidate lanes
With a provider key, OrangePro stages weak AI behaviorโsymbol links and AI-suggested candidate flows. These are review/generation worklists, not evidence:
- AI links appear as
AI-linkedsuggestions. - AI flows are stored separately from deterministic flows.
- Neither lane changes evidence tiers or denominator counts.
Use them when you want the agent to find likely service-boundary flows faster; ignore them for a deterministic-only report.
What's on the hosted platform
This repo is the free local tool. The OrangePro platform adds:
- Persistent knowledge graph across PRs and repos
- PR/CI policy gates over evidence tiers and risk deltas
- Jira / Confluence / TestRail / OpenAPI enrichment
- Cross-repo intelligence and recurring-flow memory
- Production incident correlation and regression targeting
- Team dashboards and test lifecycle management
Contributing
git clone https://github.com/OrangeproAI/orangepro-mcp.git
cd orangepro-mcp && npm ci && npm run build
npm test
PRs welcome. Please open an issue first for large changes.
MIT License ยท orangepro.ai
Sourced from the repository README.
More in AI & Agents
- PonytailMakes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.109,599
- AgentsMulti-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, and Google Antigravity39,079
- Frontend SlidesCreate beautiful slides on the web using a coding agent's frontend skills28,060
- Agent Skills Search ServerSearch and discover Agent Skills from the skills.sh registry. Powered by HAPI MCP server.24,658
- Agency Agents Zh๐ญ 267 ไธชๅณๆๅณ็จ็ AI ไธๅฎถ่ง่ฒ โ ๆฏๆ Hermes Agent/Claude Code/Cursor/Copilot ็ญ 18 ็งๅทฅๅ ท๏ผ่ฆ็ๅทฅ็จ/่ฎพ่ฎก/่ฅ้/้่็ญ 20 ไธช้จ้จใๅซ 52 ไธชไธญๅฝๅธๅบๅๅๆบ่ฝไฝ๏ผๅฐ็บขไนฆ/ๆ้ณ/ๅพฎไฟก/้ฃไนฆ/้้็ญ๏ผใๆญ้ ็ผๆๅจ agency-orchestrator๏ผไธๅฅ่ฏๅณๅฏ่ฎฉๅคไฝไธๅฎถๆ DAG ่ชๅจๅไฝใ19,868
- Watermarks RemoverStrip multi-vendor AI provenance marks: Unicode text hygiene, statistical rewrite hooks, and C2PA/metadata from PNG/JPEG/SVG/PDF/DOCX/HTML/MD17,822