optlens

Debugging LP and MILP models with a coding agent: a 400-case benchmark

Infeasible, wrong and working models of up to 837,000 constraints: what a debugging engine adds to Claude, and how six setups compare.

TL;DR

The six setups

All run Claude Sonnet 5.5. Every setup starts from the same first message: the model's context, the document itself on demand, and the question. What differs is only what the agent can call.

The model context is each constraint and variable family with its business meaning and the documented result. optlens generates it once from the model's document or code (an LLM writes the meanings; optlens checks them against the model's actual families), keeps it as JSON you can review and commit, and shows it the same way in every session, so every agent and every person describes the model in the same words. In Claude Code the plugin creates it for your model the first time you open it.

setup what the agent can call what it is
Bare Python with HiGHS, SCIP and gurobipy a coding agent on its own: the baseline
Tools the optlens tools, one call per action the MCP server, in any MCP client
Code ("optlens" in the tables) optlens preloaded in its Python, called from its own code the plugin's run_python
Claude Code + plugin tools and run_python with the optlens skill (it also offered an analyst subagent, since removed: 2.5x the cost, no gain) what an OR user installs
optchat the code setup, plus a planner prompt, an approval step and a decision log the agent for business planners
optchat + pack optchat plus a domain pack: a skill (Markdown) and helpers (Python) one pack per business domain

What the context is worth. We gave the bare agent the context in every comparison, so the tables measure the engine alone. Without it, the agent lost one case on small models and one on the largest, at 6-8 % more cost:

bare agent with the model context without it
27 small infeasible models: first fix right 26/27 25/27
cost per case $0.119 $0.129
8 hotel models of 279,000-837,000 constraints: first fix right 8/8 7/8
cost per case $0.174 $0.185

Both misses are where meanings matter. On a small anonymous model, with no meanings for its constraints, the agent proposed deleting a block of them. On a 279,000-constraint hotel model with one demand forecast typed ten times too high, it searched for the wrong value without knowing which constraints are demand and ran out of turns; with the context it found it.

What we measured

400 cases, each with a solver-verified answer key:

case type cases the question
Infeasible, one injected error (a tightened limit, an added target, a flipped sense, a tightened bound) 97 "Why is it infeasible, and how do I fix it?"
Infeasible, two independent errors or a cascading conflict (from OR-Debug-Bench) 124 the same, when one fix is not enough
Infeasible as published (Netlib, MIPLIB) 45 diagnosis only
Solves to a wrong optimum (sign, coefficient, dropped constraint, big-M; documented values off by 20-25 %) 70 "This plan looks wrong. Why?"
What-if, sensitivity, goal and why-not questions on working models 32 "What if...?"
Planner conversations (5-6 turns, edits kept between turns, a three-scenario comparison) 7 a planner's working session
Scenario comparisons 5 "How do these plans differ?"
Hotel revenue-management models of 279,000 and 837,000 constraints, infeasible or wrong 16 the same questions, on models too large to read
Planner conversations on those hotel models 2 a working session on a model too large to read
Planner conversations on a 14-night hotel model: a messy week, and a rate plan that feeds the booking plan 2 a session across two linked models

285 cases are on business models with a model document (Gurobi's modeling examples, synthetic ad-allocation and hotel revenue-management models, and OR-Debug-Bench's), 115 on anonymous Netlib and MIPLIB models.

Scoring. Every recommended fix is applied to the model and re-solved: does it solve, is the change small and within policy (no business rule deleted), does it restore the correct optimum, does it touch the actual error. What-if and conversation answers end in a JSON block checked against solver-computed keys.

Results

1. Correctness: level

First fix right, code against bare, on the same cases:

suite optlens bare
Infeasible, 27 mixed models (business and anonymous) 26/27 26/27
Infeasible, 14 business models, 2 runs 28/28 28/28
What-if, sensitivity and goal questions, 2 runs 64/64 64/64
Wrong optimum, reference objective given 19/19 19/19
Wrong optimum, no reference objective 19/19 18/19
Plausible wrong values (documented value at 80 % or 125 %) 12/12 12/12
Largest ad and yield models, wrong optimum, no reference 5/5 5/5
Two errors and cascading conflicts 40/40 40/40
Planner conversations, every turn right 7/7 7/7
Hotel models of 279,000 and 837,000 constraints 10/10 10/10
Hotel group linked by group-wide constraints, 279,000 and 837,000 constraints 6/6 6/6
Planner conversations on the linked hotel group, every turn right (9 turns) 2/2 2/2

The two are never more than one case apart. A frontier coding agent computes its own IIS, relaxes constraints and re-solves; on a model too large to read, it reads the file as text and solves a piece of it. If you evaluate LLM tools for optimization, compare them with a bare agent and a solver, not with a chat model.

2. Cost: a third to a half less, growing with model size

optlensbare
optlens, small: $0.056 per case, 5.1 turns, 80/81 right$0.056bare, small: $0.085 per case, 6.2 turns, 79/81 right$0.085−35%smalloptlens, mid: $0.066 per case, 5.7 turns, 56/56 right$0.066bare, mid: $0.111 per case, 8.3 turns, 56/56 right$0.111−40%midoptlens, large: $0.050 per case, 3.7 turns, 13/13 right$0.050bare, large: $0.087 per case, 6.7 turns, 13/13 right$0.087−42%largeoptlens, too large to read: $0.069 per case, 6.1 turns, 16/16 right$0.069bare, too large to read: $0.133 per case, 10.0 turns, 16/16 right$0.133−48%too large to read
Mean cost per case and run, same cases; the percentage is optlens's saving. Hover a bar for turns and correctness.
model size right (optlens / bare) optlens: cost, turns bare: cost, turns optlens saves
small (constraints × variables < 10⁴) 80/81 / 79/81 $0.056, 5.1 $0.085, 6.2 35 %
mid (< 10⁶) 56/56 / 56/56 $0.066, 5.7 $0.111, 8.3 40 %
large (≥ 10⁶) 13/13 / 13/13 $0.050, 3.7 $0.087, 6.7 42 %
too large to read (50,000+ constraints) 16/16 / 16/16 $0.069, 6.1 $0.133, 10.0 48 %

Mean per case and run, on the same cases. The bare agent gets there, but it rebuilds the diagnostic machinery in a scratch script every session.

3. Time on models too large to read

Hotel revenue-management models at 4 hotels × 365 nights (279,000 constraints, a 176 MB file) and 12 hotels × 365 nights (837,000 constraints, 530 MB), with injected errors; a second set links the hotels with group-wide constraints, so the model no longer splits into one hotel at a time.

optlensbare
hotel models (10)optlens, hotel models (10): 26 min of tool time26 minbare, hotel models (10): 40 min of tool time40 min−35%linked hotel group (6)optlens, linked hotel group (6): 25 min of tool time25 minbare, linked hotel group (6): 64 min of tool time64 min−61%conversations, linked group (2)optlens, conversations, linked group (2): 20 min of tool time20 minbare, conversations, linked group (2): 46 min of tool time46 min−56%
Total tool time per suite (models of 279,000 and 837,000 constraints); the percentage is optlens's saving.
optlens: tool time, cost bare: tool time, cost right (optlens / bare)
hotel models (10 cases) 26 min, $0.73 40 min, $1.40 10/10 / 10/10
linked hotel group (6 cases) 25 min, $0.53 64 min, $0.86 6/6 / 6/6
planner conversations on the linked group (2, 9 turns) 20 min, $0.40 46 min, $0.60 9/9 / 9/9

optlens finds the conflict from the solver's own proof of infeasibility, so it never needs the whole model: 13 to 37 conflicting constraints out of 279,000-837,000, found in 4 to 58 seconds.

4. What you can ask

Every capability is a tool in Python, the MCP server and the plugin:

you want to know optlens tools tested in
what the model is and means get_model_overview, query_constraints, query_variables, read_model_document, save_model_context every suite (the context)
why it is infeasible compute_iis, feasibility_relaxation, drop_test the infeasible suites
what fixes it fix_menu, suspicious_values (with each flag's undo solved) the infeasible and wrong-optimum suites
what happens if... modify_and_resolve, try_options, version_history what-if questions, planner conversations
how two plans differ compare_versions (against the closest optimal plan) what-if questions, planner conversations
what a limit is worth, how far it can go marginal_value, sensitivity_report, attainable_limit what-if and goal questions
why the plan doesn't do X why_not why-not questions
how several models or scenarios compare add_model, compare_models scenario comparisons, the three-scenario conversation
anything else, in code run_python (the engine preloaded as session) the code setup, every suite

5. Tools, code or plugin

14 business infeasible models and 32 what-if questions right cost per case
bare 46/46 $0.061
tools (one call per action) 45/46 $0.052
code (optlens in the agent's Python) 46/46 $0.043

Calling the engine from code is the cheapest setup: several steps per call, one model load. Claude Code with the plugin answered 5/5 questions comparing three ad-allocation scenarios (three 9,680-variable MIPs opened side by side: status and objectives, a what-if on one scenario, the smallest feasible budget of another), at $0.14 a question (a small test).

6. optchat and domain packs

Turns answered right; mean cost per conversation:

conversation code optchat optchat + hotel pack
Everyday planner conversations (7) 40/40, $0.22 40/40, $0.28 no pack for these models
Linked hotel group, 279,000-837,000 constraints (2) 9/9, $0.20 9/9, $0.38 9/9, $0.41
Messy week on the hotel model (1) 5/5, $0.23 5/5, $0.41 15/15 over 3 runs, $0.36
Two-stage rate plan feeding the booking plan (1) 1/4, $0.39 2/4, $0.68 4/4, $0.19

On a single model, optchat answers as correctly as the code setup at a modest extra cost: its planner prompt asks for checks, plain language and the planner's approval before a change becomes the plan. Where models are linked, the pack decides: when the planner changes the rate plan, the pack's helper carries the new prices and demand into the booking plan exactly, in 1.5 seconds of tool time. Without it, both agents rebuilt or estimated the link (110 and 277 seconds) and got turns wrong.

A new business domain needs no code change to the engine or the agents:

you add format what it gives
the model and its document LP, MPS or Python; Markdown every question above (all the Gurobi example models ran this way)
a pack's skill Markdown the domain's levers and words in every answer
a pack's helpers (optional) Python exact steps across linked models, data checks against the document

7. Plan comparison on large MIPs

A MIP with many equally good plans re-solves to any of them, so a plain diff mostly shows the solver's choice among ties: over 99 % of the reported changes on the ad min-cost what-ifs. optlens compares against the optimal plan closest to the old one (one extra solve, ≤ 1.5 s here):

ad min-cost what-if variables reported as changed
plain diff 755
optlens (closest optimal plan) 9

In a planner conversation, "campaign 3 now needs 48 impressions" is answered with the four slot changes the edit forces, adding up to the +32.4 cost, instead of a 779-variable reshuffle.

8. Checked fixes for suspicious values

optlens flags inputs that break their own pattern (a limit unlike its siblings', a flipped sense, an out-of-scale coefficient, a missing constraint) and solves each flag's undo, alone and together. On 221 cases with a known error, it hands over a verified correct fix for 47, and a wrong one for 2.

What the engine adds when the agent is already right

Cost and speed 35-48 % less spend, fewer turns, 35-61 % less time on models too large to read
The same facts every time conflicts, flags, fixes and plan diffs come from tested code, not a script written this session
Every number checked recommended fixes are re-solved before they are offered
Any solver, local HiGHS and SCIP out of the box, your own Gurobi licence; a hard time limit on every solve; LP, MPS, Pyomo, gurobipy and PuLP

Limits

Try it

pip install "optlens[scip,mcp] @ git+https://github.com/jjd-lab/optlens"

Then connect it once: in Claude Code, claude plugin install optlens --marketplace jjd-lab/optlens; in Codex, codex mcp add optlens -- optlens-mcp; in Cursor, VS Code or Claude Desktop, one line in the MCP config (the README has each). The agent gets the tools and the method with them. optchat and domain packs for your models: hello@optlens.dev.

The benchmark itself stays private, so that it keeps measuring new versions of the engine fairly; the tables above are generated from its recorded runs. Issues with your own models are the most useful feedback we can get: open a GitHub issue or write to hello@optlens.dev.

Credits

Business example models from Gurobi's modeling-examples (Apache-2.0). Two-error and cascading cases from OR-Debug-Bench / ORLoopBench (Ao, Simchi-Levi and Wang, arXiv 2601.21008; data CC BY 4.0), with constraint names changed and cases filtered. Netlib LP (infeasible set collected by J. Chinneck) and MIPLIB 2017 instances.