DEV API LAB / FIELD GUIDE 03

Evaluate LLM APIs with evidence, not demos

Build a task-specific evaluation set, investigate failure patterns, and make a reproducible release decision.

By DevAPI.com™7 min readLLM Directory
Evaluate LLM APIs with evidence, not demos — neon typographic artwork with DevAPI.com™ branding

A language-model integration can look excellent in a demonstration and disappoint on ordinary work. The demonstration usually contains a clear question, a convenient source document, and a person who already knows what a good answer looks like. Production traffic contains ambiguity, missing context, outdated documents, and requests that the feature should decline to answer. Evaluating those conditions is a design task, not a final checkbox before release.

This guide proposes an evaluation process for a documentation assistant. The assistant answers questions from an approved set of product documents and should acknowledge when those documents do not support an answer. The process is useful for the LLM Dev API topic without pretending that one public benchmark identifies the best model for every private workflow.

Define the decision the evaluation supports

Before collecting examples, decide what you are trying to approve. Is the team choosing between two configurations, checking a new retrieval method, or determining whether a feature is ready for a limited rollout? A vague goal such as improve quality makes it easy to collect impressive examples without producing a useful decision.

Write the acceptance criteria in product language. For the documentation assistant, a useful answer should address the question, rely on the selected documents, identify relevant evidence, and avoid inventing unsupported procedures. A correct refusal is valuable when the source material is insufficient. That last point prevents an evaluation from rewarding an assistant simply for always producing an answer.

Use an explicit evaluation framework

The OpenAI evaluation best-practices guide recommends defining an objective, collecting relevant data, choosing metrics, comparing runs, and evaluating over time. Use that structure as a starting point, not as a replacement for your own acceptance criteria. A metric only matters when it helps explain whether the feature is useful for its intended task.

Keep the experiment configuration recorded alongside the results. Include the model identifier, prompt revision, retrieval settings, document snapshot, and any output schema. Without those inputs, an apparent improvement can be difficult to reproduce. It may reflect a changed source document rather than a better model or prompt.

Build a representative case collection

Start with the actual tasks the assistant is meant to support. Include straightforward lookups, questions requiring several sections, ambiguous terminology, outdated assumptions, and questions outside the allowed source scope. Use a deliberate collection rather than only the requests that were easy to write. The difficult cases reveal where the product boundary needs clarification.

For each case, preserve the question and the exact source revision. Write a short explanation of what a satisfactory answer must contain and what it must not claim. A reference answer can help, but do not require the model to reproduce one author’s wording. The evaluation should reward correct meaning and evidence, not superficial similarity to a preferred sentence.

Separate development cases from held-out cases

Use one collection to improve the system and another to check whether improvements generalize. Repeatedly editing the prompt against the same small set can make that set less informative. The feature may learn the shape of the examples through your iteration process while still failing on a fresh question with the same underlying difficulty.

Do not turn the held-out collection into a secret leaderboard that nobody can investigate. When a case fails, preserve it and analyze the failure. Then decide whether to move it into the development collection and replace it with another independently constructed case. Document this maintenance process so the evaluation remains both useful for learning and meaningful for release decisions.

Score different dimensions separately

A single number can hide important disagreements. An answer may be factually correct but fail to address the question. Another may be responsive but invent a setting that does not exist. Score answer relevance, evidence support, completeness, and appropriate abstention separately. Add operational dimensions such as latency and observed usage without pretending they are the same kind of quality.

Define the scoring rubric before comparing configurations. For evidence support, a reviewer might distinguish fully supported, partly supported, and unsupported claims. For completeness, identify the required steps or concepts for that specific question. Avoid a rubric whose labels are only excellent, good, and bad; reviewers need concrete reasons to apply those labels consistently.

Diagnose retrieval and generation independently

When the answer is wrong, first check what evidence the assistant received. The model cannot reliably cite a paragraph that the retrieval stage never supplied. Conversely, a correct paragraph in context does not guarantee that the generated answer uses it correctly. Preserve the retrieved document identifiers and relevant passages for the evaluation run.

Create a diagnostic mode that supplies known relevant context directly. Comparing that mode with the full retrieval pipeline helps locate the failure. If both fail, the issue may involve interpretation or the output instructions. If only the full pipeline fails, inspect indexing, filtering, document freshness, and query construction before spending time polishing the final response prompt.

Review disagreements rather than averaging them away

Have reviewers explain uncertain or disputed scores. A disagreement can reveal an ambiguous product requirement, an incomplete reference answer, or an unclear rubric. Resolve the underlying question rather than simply averaging the ratings into a more precise-looking number. The evaluation is helping the team define the intended behavior as well as measure it.

A model-based grader can assist with comparison, but it should not become an unquestioned authority over the feature being evaluated. Check grader decisions against human judgments on representative cases. Keep the grader’s configuration versioned, and investigate whether it systematically favors verbosity, a particular answer style, or unsupported confidence.

Measure operational behavior honestly

Record observed request duration and usage with the conditions of the run. A local experiment on a quiet network is not a universal latency guarantee. Compare configurations under similar conditions, and distinguish time to first visible output from time to a complete validated answer. An attractive streaming experience can still end in an incomplete or invalid result.

For a cost worksheet, use actual observed units and the applicable provider pricing at implementation time. Do not publish remembered prices as permanent facts. Include repeated calls, retrieval work, and unsuccessful attempts when they contribute to the feature’s operating cost. A cheap successful response is not the entire cost of serving a difficult task reliably.

Test the boundary cases deliberately

Add cases in which the source material contains instructions that conflict with the application’s task. Add a document that looks authoritative but is outside the approved collection. Test a question that requests confidential data the assistant should not have. The expected result may be refusal, clarification, or a bounded explanation of the available source scope.

Also test interrupted and resumed interactions for a chat AI interface. A question that is answerable in isolation may become ambiguous after earlier conversation turns. Keep account identity and permissions outside the model context. Evaluation can reveal a boundary problem, but deterministic authorization checks must still enforce that boundary in the application.

Turn findings into release criteria

Summarize results by task category and failure type, not only by an overall average. Identify which failures block release and which require interface changes or narrower scope. A limited rollout can be a sensible next step when the system performs well on a specific task but is not ready for a broader promise.

Record the decision with the tested configuration and known limitations. If a model, prompt, document pipeline, or tool policy changes, decide which cases must run again. The safe tool-calling guide explains why a feature that performs actions needs additional tests of approval, authorization, and duplicate protection beyond the quality of its generated text.

Conclusion: evaluate the product you are building

A good evaluation makes uncertainty visible and turns it into concrete engineering work. It does not manufacture a universal ranking from a narrow experiment. For a documentation assistant, the key question is whether the system answers the intended questions with appropriate evidence and restraint. Keep the cases representative, the configuration reproducible, and the release decision tied to the actual task rather than the most impressive demonstration.