Home

Passing the tests doesn't mean the model is right

Have a large language model turn a scheduling problem described in text into an optimization model, then check whether it actually got it right. A program that passes the tests can be wrong on a different set of numbers.

TYPE
Evaluating automatic modeling by large language models
STATUS
Working paper
WHEN
2026

The problem

Scheduling problems (which truck goes first, which machine does which job) usually require someone who knows operations research to write them up as an optimization model: what the variables are, what the constraints are, what the objective is, and then hand it to a solver. These days you can have a large language model read a text description and write the model directly.

The trouble is how to know it got it right. The usual way is to test it on a set of cases: if the answers match the reference answers, it passes. But matching only means it was right on that set of numbers. Say the model is missing a constraint, and on this set of numbers that constraint happens not to matter, the answer still comes out right. Switch to a different set of numbers and it's wrong (fig. 1).

Passing the tests and building the right model are different thingsA paragraphdescribing a scheduling problemLLMreads it, models itOptimization modelvariables, constraints, objectiveSolveThe numbers in the testNew numbers, same problemanswer matcheswrong answer

FIG. 1 · The same model, two sets of numbers

A sketch.

Where it stands