Visual question answering models are graded on datasets built by someone else, asking questions a research team wrote, about images a research team chose. None of those three things are your users, your questions, or your photos. A strong benchmark number is a hypothesis about how well that gap doesn't matter — not a guarantee.

The gap shows up in predictable places. A model that scores well on crisp, well-lit, single-subject benchmark images will wobble on a blurry warehouse photo taken at an angle under fluorescent light, because nothing in training or eval ever looked like that. The benchmark wasn't wrong. It was just answering a different, easier question than the one your product actually asks.

What's worked for me isn't more benchmarking — it's smaller, uglier, real evaluation:

  • Pull 50 real examples before you pull one more public dataset. Actual inputs from the actual product, even a rough sample, tell you more about where a model breaks than another point of leaderboard accuracy.
  • Write down what "correct" means for your use case, not the dataset's. A VQA model answering "what color is this" for an accessibility tool has a much stricter bar than one answering it for a trivia game, even on the identical image.
  • Re-run the eval set every time the model, prompt, or preprocessing changes. Evaluation debt accumulates exactly like technical debt — quietly, until a change upstream breaks something downstream that nobody was watching.

The benchmark number is still useful — it's just useful for comparing models to each other, not for predicting how one will behave on your users' actual, messy input. Those are different questions, and conflating them is where evaluation debt comes from.