Passing the tests doesn't mean the model is right
Have a large language model turn a scheduling problem described in text into an optimization model, then check whether it actually got it right. A program that passes the tests can be wrong on a different set of numbers.
- TYPE
- Evaluating automatic modeling by large language models
- STATUS
- Working paper
- WHEN
- 2026
The problem
Scheduling problems (which truck goes first, which machine does which job) usually require someone who knows operations research to write them up as an optimization model: what the variables are, what the constraints are, what the objective is, and then hand it to a solver. These days you can have a large language model read a text description and write the model directly.
The trouble is how to know it got it right. The usual way is to test it on a set of cases: if the answers match the reference answers, it passes. But matching only means it was right on that set of numbers. Say the model is missing a constraint, and on this set of numbers that constraint happens not to matter, the answer still comes out right. Switch to a different set of numbers and it's wrong (fig. 1).
FIG. 1 · The same model, two sets of numbers
A sketch.
Where it stands
- This is a paper still being written. How to check, and how much the checking turns up, will go up once it is published.