Debugging LP and MILP models with a coding agent: a 400-case benchmark
Infeasible, wrong and working models of up to 837,000 constraints: what a debugging engine adds to Claude, and how six setups compare.
TL;DR
- A coding agent with a solver already gets the answers right: 95-100 % in every suite, with or without the engine.
- The engine's edge is cost and speed: the same correct answers at 35-48 % less cost, and on models too large to read, 35-61 % less time.
- On linked models, a domain pack wins on correctness too: where a rate-plan change must flow into the booking plan, optchat with the hotel pack answered 4/4 turns right; the same agent without the pack 2/4, and the coding agent with the engine 1/4.
- Plan comparisons show what an edit forces: 755 "changed" variables become the 9 the edit actually forces.
- The engine, its MCP server and the Claude Code plugin are open source: github.com/jjd-lab/optlens. optchat, the planner agent, is available on request: hello@optlens.dev.
The six setups
All run Claude Sonnet 5.5. Every setup starts from the same first message: the model's context, the document itself on demand, and the question. What differs is only what the agent can call.
The model context is each constraint and variable family with its business meaning and the documented result. optlens generates it once from the model's document or code (an LLM writes the meanings; optlens checks them against the model's actual families), keeps it as JSON you can review and commit, and shows it the same way in every session, so every agent and every person describes the model in the same words. In Claude Code the plugin creates it for your model the first time you open it.
| setup | what the agent can call | what it is |
|---|---|---|
| Bare | Python with HiGHS, SCIP and gurobipy | a coding agent on its own: the baseline |
| Tools | the optlens tools, one call per action | the MCP server, in any MCP client |
| Code ("optlens" in the tables) | optlens preloaded in its Python, called from its own code | the plugin's run_python |
| Claude Code + plugin | tools and run_python with the optlens skill (it also offered an analyst subagent, since removed: 2.5x the cost, no gain) |
what an OR user installs |
| optchat | the code setup, plus a planner prompt, an approval step and a decision log | the agent for business planners |
| optchat + pack | optchat plus a domain pack: a skill (Markdown) and helpers (Python) | one pack per business domain |
What the context is worth. We gave the bare agent the context in every comparison, so the tables measure the engine alone. Without it, the agent lost one case on small models and one on the largest, at 6-8 % more cost:
| bare agent | with the model context | without it |
|---|---|---|
| 27 small infeasible models: first fix right | 26/27 | 25/27 |
| cost per case | $0.119 | $0.129 |
| 8 hotel models of 279,000-837,000 constraints: first fix right | 8/8 | 7/8 |
| cost per case | $0.174 | $0.185 |
Both misses are where meanings matter. On a small anonymous model, with no meanings for its constraints, the agent proposed deleting a block of them. On a 279,000-constraint hotel model with one demand forecast typed ten times too high, it searched for the wrong value without knowing which constraints are demand and ran out of turns; with the context it found it.
What we measured
400 cases, each with a solver-verified answer key:
| case type | cases | the question |
|---|---|---|
| Infeasible, one injected error (a tightened limit, an added target, a flipped sense, a tightened bound) | 97 | "Why is it infeasible, and how do I fix it?" |
| Infeasible, two independent errors or a cascading conflict (from OR-Debug-Bench) | 124 | the same, when one fix is not enough |
| Infeasible as published (Netlib, MIPLIB) | 45 | diagnosis only |
| Solves to a wrong optimum (sign, coefficient, dropped constraint, big-M; documented values off by 20-25 %) | 70 | "This plan looks wrong. Why?" |
| What-if, sensitivity, goal and why-not questions on working models | 32 | "What if...?" |
| Planner conversations (5-6 turns, edits kept between turns, a three-scenario comparison) | 7 | a planner's working session |
| Scenario comparisons | 5 | "How do these plans differ?" |
| Hotel revenue-management models of 279,000 and 837,000 constraints, infeasible or wrong | 16 | the same questions, on models too large to read |
| Planner conversations on those hotel models | 2 | a working session on a model too large to read |
| Planner conversations on a 14-night hotel model: a messy week, and a rate plan that feeds the booking plan | 2 | a session across two linked models |
285 cases are on business models with a model document (Gurobi's modeling examples, synthetic ad-allocation and hotel revenue-management models, and OR-Debug-Bench's), 115 on anonymous Netlib and MIPLIB models.
Scoring. Every recommended fix is applied to the model and re-solved: does it solve, is the change small and within policy (no business rule deleted), does it restore the correct optimum, does it touch the actual error. What-if and conversation answers end in a JSON block checked against solver-computed keys.
Results
1. Correctness: level
First fix right, code against bare, on the same cases:
| suite | optlens | bare |
|---|---|---|
| Infeasible, 27 mixed models (business and anonymous) | 26/27 | 26/27 |
| Infeasible, 14 business models, 2 runs | 28/28 | 28/28 |
| What-if, sensitivity and goal questions, 2 runs | 64/64 | 64/64 |
| Wrong optimum, reference objective given | 19/19 | 19/19 |
| Wrong optimum, no reference objective | 19/19 | 18/19 |
| Plausible wrong values (documented value at 80 % or 125 %) | 12/12 | 12/12 |
| Largest ad and yield models, wrong optimum, no reference | 5/5 | 5/5 |
| Two errors and cascading conflicts | 40/40 | 40/40 |
| Planner conversations, every turn right | 7/7 | 7/7 |
| Hotel models of 279,000 and 837,000 constraints | 10/10 | 10/10 |
| Hotel group linked by group-wide constraints, 279,000 and 837,000 constraints | 6/6 | 6/6 |
| Planner conversations on the linked hotel group, every turn right (9 turns) | 2/2 | 2/2 |
The two are never more than one case apart. A frontier coding agent computes its own IIS, relaxes constraints and re-solves; on a model too large to read, it reads the file as text and solves a piece of it. If you evaluate LLM tools for optimization, compare them with a bare agent and a solver, not with a chat model.
2. Cost: a third to a half less, growing with model size
| model size | right (optlens / bare) | optlens: cost, turns | bare: cost, turns | optlens saves |
|---|---|---|---|---|
| small (constraints × variables < 10⁴) | 80/81 / 79/81 | $0.056, 5.1 | $0.085, 6.2 | 35 % |
| mid (< 10⁶) | 56/56 / 56/56 | $0.066, 5.7 | $0.111, 8.3 | 40 % |
| large (≥ 10⁶) | 13/13 / 13/13 | $0.050, 3.7 | $0.087, 6.7 | 42 % |
| too large to read (50,000+ constraints) | 16/16 / 16/16 | $0.069, 6.1 | $0.133, 10.0 | 48 % |
Mean per case and run, on the same cases. The bare agent gets there, but it rebuilds the diagnostic machinery in a scratch script every session.
3. Time on models too large to read
Hotel revenue-management models at 4 hotels × 365 nights (279,000 constraints, a 176 MB file) and 12 hotels × 365 nights (837,000 constraints, 530 MB), with injected errors; a second set links the hotels with group-wide constraints, so the model no longer splits into one hotel at a time.
| optlens: tool time, cost | bare: tool time, cost | right (optlens / bare) | |
|---|---|---|---|
| hotel models (10 cases) | 26 min, $0.73 | 40 min, $1.40 | 10/10 / 10/10 |
| linked hotel group (6 cases) | 25 min, $0.53 | 64 min, $0.86 | 6/6 / 6/6 |
| planner conversations on the linked group (2, 9 turns) | 20 min, $0.40 | 46 min, $0.60 | 9/9 / 9/9 |
optlens finds the conflict from the solver's own proof of infeasibility, so it never needs the whole model: 13 to 37 conflicting constraints out of 279,000-837,000, found in 4 to 58 seconds.
4. What you can ask
Every capability is a tool in Python, the MCP server and the plugin:
| you want to know | optlens tools | tested in |
|---|---|---|
| what the model is and means | get_model_overview, query_constraints, query_variables, read_model_document, save_model_context |
every suite (the context) |
| why it is infeasible | compute_iis, feasibility_relaxation, drop_test |
the infeasible suites |
| what fixes it | fix_menu, suspicious_values (with each flag's undo solved) |
the infeasible and wrong-optimum suites |
| what happens if... | modify_and_resolve, try_options, version_history |
what-if questions, planner conversations |
| how two plans differ | compare_versions (against the closest optimal plan) |
what-if questions, planner conversations |
| what a limit is worth, how far it can go | marginal_value, sensitivity_report, attainable_limit |
what-if and goal questions |
| why the plan doesn't do X | why_not |
why-not questions |
| how several models or scenarios compare | add_model, compare_models |
scenario comparisons, the three-scenario conversation |
| anything else, in code | run_python (the engine preloaded as session) |
the code setup, every suite |
5. Tools, code or plugin
| 14 business infeasible models and 32 what-if questions | right | cost per case |
|---|---|---|
| bare | 46/46 | $0.061 |
| tools (one call per action) | 45/46 | $0.052 |
| code (optlens in the agent's Python) | 46/46 | $0.043 |
Calling the engine from code is the cheapest setup: several steps per call, one model load. Claude Code with the plugin answered 5/5 questions comparing three ad-allocation scenarios (three 9,680-variable MIPs opened side by side: status and objectives, a what-if on one scenario, the smallest feasible budget of another), at $0.14 a question (a small test).
6. optchat and domain packs
Turns answered right; mean cost per conversation:
| conversation | code | optchat | optchat + hotel pack |
|---|---|---|---|
| Everyday planner conversations (7) | 40/40, $0.22 | 40/40, $0.28 | no pack for these models |
| Linked hotel group, 279,000-837,000 constraints (2) | 9/9, $0.20 | 9/9, $0.38 | 9/9, $0.41 |
| Messy week on the hotel model (1) | 5/5, $0.23 | 5/5, $0.41 | 15/15 over 3 runs, $0.36 |
| Two-stage rate plan feeding the booking plan (1) | 1/4, $0.39 | 2/4, $0.68 | 4/4, $0.19 |
On a single model, optchat answers as correctly as the code setup at a modest extra cost: its planner prompt asks for checks, plain language and the planner's approval before a change becomes the plan. Where models are linked, the pack decides: when the planner changes the rate plan, the pack's helper carries the new prices and demand into the booking plan exactly, in 1.5 seconds of tool time. Without it, both agents rebuilt or estimated the link (110 and 277 seconds) and got turns wrong.
A new business domain needs no code change to the engine or the agents:
| you add | format | what it gives |
|---|---|---|
| the model and its document | LP, MPS or Python; Markdown | every question above (all the Gurobi example models ran this way) |
| a pack's skill | Markdown | the domain's levers and words in every answer |
| a pack's helpers (optional) | Python | exact steps across linked models, data checks against the document |
7. Plan comparison on large MIPs
A MIP with many equally good plans re-solves to any of them, so a plain diff mostly shows the solver's choice among ties: over 99 % of the reported changes on the ad min-cost what-ifs. optlens compares against the optimal plan closest to the old one (one extra solve, ≤ 1.5 s here):
| ad min-cost what-if | variables reported as changed |
|---|---|
| plain diff | 755 |
| optlens (closest optimal plan) | 9 |
In a planner conversation, "campaign 3 now needs 48 impressions" is answered with the four slot changes the edit forces, adding up to the +32.4 cost, instead of a 779-variable reshuffle.
8. Checked fixes for suspicious values
optlens flags inputs that break their own pattern (a limit unlike its siblings', a flipped sense, an out-of-scale coefficient, a missing constraint) and solves each flag's undo, alone and together. On 221 cases with a known error, it hands over a verified correct fix for 47, and a wrong one for 2.
What the engine adds when the agent is already right
| Cost and speed | 35-48 % less spend, fewer turns, 35-61 % less time on models too large to read |
| The same facts every time | conflicts, flags, fixes and plan diffs come from tested code, not a script written this session |
| Every number checked | recommended fixes are re-solved before they are offered |
| Any solver, local | HiGHS and SCIP out of the box, your own Gurobi licence; a hard time limit on every solve; LP, MPS, Pyomo, gurobipy and PuLP |
Limits
- One model family (Claude Sonnet 5.5); one or two runs per setting (the messy week with the pack: three).
- Every correctness suite is saturated: these cases cannot show further gains in correctness.
- The largest models are one synthetic family (hotel revenue management): 16 cases and 2 conversations.
- The plugin row is 5 questions; the optchat rows are one run each, the two-stage rate plan one conversation.
- Gurobi was tested only with its size-limited licence.
Try it
pip install "optlens[scip,mcp] @ git+https://github.com/jjd-lab/optlens"
Then connect it once: in Claude Code, claude plugin install optlens --marketplace jjd-lab/optlens; in Codex,
codex mcp add optlens -- optlens-mcp; in Cursor, VS Code or Claude Desktop, one line in the MCP config (the README
has each). The agent gets the tools and the method with them. optchat and domain packs for your models:
hello@optlens.dev.
The benchmark itself stays private, so that it keeps measuring new versions of the engine fairly; the tables above are generated from its recorded runs. Issues with your own models are the most useful feedback we can get: open a GitHub issue or write to hello@optlens.dev.
Credits
Business example models from Gurobi's modeling-examples (Apache-2.0). Two-error and cascading cases from OR-Debug-Bench / ORLoopBench (Ao, Simchi-Levi and Wang, arXiv 2601.21008; data CC BY 4.0), with constraint names changed and cases filtered. Netlib LP (infeasible set collected by J. Chinneck) and MIPLIB 2017 instances.