A request in Chinese for a simple tomato-and-egg recipe produced an unexpectedly elaborate response from Grok 4.5: ingredient ranges, optional tomato ketchup, long cooking tips—and two links to posts on X. Yet the recorded call says web search was neither requested nor effective. That does not establish anything about the recipe’s accuracy, but it is a vivid example of the problem examined here: when a user’s real requirement is a particular presentation, a polished answer can bring along a lot more than was asked for.

Disclosure: I work with OrcaRouter and used it to run this evaluation. One OpenAI-compatible key gave me access to every model in this test, without changing how any model answered.

Compare the models in this article through OrcaRouter’s model catalog.

This is not a universal model ranking. It is a small case study of how eight models presented answers to two short custom prompts: one asking how to make 西红柿炒蛋 (tomato-and-egg stir-fry), and one asking for a casual, text-message-like science-fiction recommendation. The run contained 16 selected API calls—eight models times two questions—with one selected answer per model-question cell.

That is an observation set, not proof of stable model behavior.

The format was the practical test

Neither prompt demanded a formal JSON schema, an exact word count, or a rigid list structure. Instead, each set a softer formatting expectation: a useful answer to a simple cooking question, and a recommendation that felt conversational rather than like a catalog entry. That distinction matters. The evidence can show how these selected outputs looked; it cannot establish broad “format-compliance” rates for the models.

Here is the compact picture of the selected calls:

PromptWhat appeared in the eight selected responsesWhat that does—and does not—show
Chinese recipe requestAll eight answers expanded into structured recipes with headings, ingredients, steps, and tips.The models generally chose a tutorial-like format in this one round; it does not show that the recipes were correct or that they would always be similarly detailed.
Casual sci-fi requestAll eight recommended Project Hail Mary; the answers also used informal language and ended or opened with questions about the reader’s preferences.The selected outputs converged strongly on one recommendation and a chatty posture; it does not validate the book descriptions or prove a lasting stylistic tendency.

The recipe responses reveal the cost of “helpfulness” when brevity or a clean template may be what the user actually needs. GPT-5.6 Sol gave a relatively compact five-step recipe with an ingredient list and a short tips section. GPT-5.6 Terra added multiple optional ingredients, detailed timing, and three extra preference-based variants. Gemini 3.5 Flash supplied a much longer “zero-failure” tutorial with separate sections for preparation, egg frying, tomato frying, combination, and advanced tips.

None of those choices is inherently wrong. But they are different answers to a hidden question: should the assistant provide the minimum usable response, or anticipate every possible preference? For a reader pasting text into a recipe card, a shopping list, or a narrow app field, that difference is often more consequential than the underlying prose quality.

The casual-recommendation prompt showed a different kind of convergence. GPT-5.6 Terra offered one main recommendation, one alternative, and a short question about the reader’s mood. Claude Fable 5 gave a looser, more expansive set of suggestions, complete with “always down to talk sci-fi” and “Extremely relatable honestly lol.” GLM-5.2 opened with “Dude yes” and urged the reader to “drop everything and get” the same book.

These outputs broadly fit the requested conversational setting in this single sample. But a casual voice is not the same as a controlled output format. Emojis, enthusiasm, plot setup, extra recommendations, and follow-up questions may feel friendly—or may be unwanted additions if the user needed one title in one sentence.

What was observed, and what was not scored

All 16 selected calls completed with an ok status. Search was recorded as not requested and not effective for every selected call.

The supplied material does not provide a valid judge-quality score for these cells, so there is no scored basis here for declaring a best model. Old or failed judge versions would be audit-trail material rather than quality ground truth; none should be used to manufacture a ranking.

Likewise, the recorded response times are descriptive, not a speed league table. In these selected calls, Gemini 3.5 Flash recorded 2.6 seconds for the conversational prompt and 6.99 seconds for the recipe prompt; GLM-5.2 recorded 11.38 and 21.69 seconds respectively. A single call can be affected by conditions this small test does not isolate, so those numbers should not be read as general superiority claims.

The same restraint applies to observed billing fields and token counts. They describe these particular gateway/API calls, not consumer subscription products, and they cannot tell us whether vendor effort labels would represent comparable compute across providers.

Why a famous benchmark does not settle this question

Humanity’s Last Exam is a primary-source academic benchmark for closed-ended questions, but it is not a format-compliance benchmark and cannot rank these two custom prompts. That is not a flaw in that benchmark; it is a reminder that an evaluation only answers the question it was designed to ask.

For formatting-sensitive work, the useful question is narrower: can a model follow your contract—your fields, length cap, language, tone, exclusions, and order—repeatedly?

Practical takeaways

  • If format is the deliverable, write the format explicitly. “Conversational” invited different levels of informality and detail across the selected answers.
  • Test the exact template you plan to deploy, including what must not appear: extra citations, headings, follow-up questions, or optional variations.
  • Run repeated samples. One successful-looking response is only one observation.
  • Separate formatting checks from factual checks. A valid-looking format is not evidence that an answer is factually correct.

Limitations

This study used two custom prompts and one selected answer per model-question cell. It covers 16 selected gateway/API calls from one approved run, not consumer-product behavior, broad model reliability, factual correctness, or an overall capability ranking. The results are therefore best read as concrete examples of presentation choices under light formatting pressure—not a verdict on which model is “best.”

Explore the Models

Explore the current catalog on OrcaRouter Models.

This evaluation was run through OrcaRouter. The author works with OrcaRouter; model access does not imply affiliation with, endorsement by, or sponsorship from model providers.

Model names and logos are used descriptively. All trademarks belong to their respective owners.

Sources