LLM Models, Benchmarks and Evals: How to Choose a Model on Evidence, Not Leaderboards
How large language models differ, what the popular benchmarks really measure — and where they mislead — and how to build your own evals so you pick and change models with confidence.
By RudraAI · · 12 min read

Key takeaways
- Benchmarks are useful for a shortlist; your own evals should make the final decision.
- Watch for contamination, saturation and different test harnesses when comparing leaderboard scores.
- A good eval set is 50–200 real examples with code-based checks first and an LLM judge for the rest.
- Pick the cheapest, fastest model that clears your quality bar — then route harder requests to bigger models.
On this page
- Benchmarks vs evals: the short answer
- What an LLM is, in four ideas
- The model landscape
- What actually differs between models
- Estimating cost: a worked example
- Benchmarks: what they measure
- How to read a leaderboard without being misled
- Evals: your own benchmark
- Measuring hallucinations
- Evals for RAG and agents
- Tools that help
- Prompting, RAG or fine-tuning?
- A practical model-selection recipe
- LLM glossary
New models ship almost every month, each with a chart showing it beating the last one. If you're building anything on top of LLMs, you need a way to cut through that: which model is actually best for your task, at a price and speed you can live with — and how will you know if switching models breaks something?
This guide covers three things: how LLMs differ, what public benchmarks measure (and don't), and how to build your own evals — the tests that turn model choice from guesswork into engineering.
Benchmarks vs evals: the short answer
An LLM benchmark is a standard public test used to compare models on general abilities; an LLM eval is your own repeatable test that measures how well your specific application performs on your specific task. Benchmarks answer "which models are strong in general?" Evals answer "which model, prompt and setup work best for us?" You need both: benchmarks to shortlist, evals to decide.
What an LLM is, in four ideas
- Tokens. Models read and write text as tokens — word pieces of roughly three-quarters of a word in English. Pricing, speed and limits are all measured in tokens.
- Next-token prediction. A transformer network predicts the most likely next token, over and over. Everything else — answering, coding, reasoning — emerges from doing this extremely well.
- Parameters. The learned weights of the network. More parameters generally means more capability and more cost, though training data and techniques matter as much as size.
- Context window. How many tokens the model can consider at once — your instructions, documents, conversation and its own answer. Large windows are useful, but models don't use every part of a huge context equally well.
Modern models are built in stages: pre-training on vast amounts of text, instruction tuning to follow requests, preference training (such as RLHF) to be helpful and safe, and increasingly reasoning training so the model can work through a problem step by step before answering.
The model landscape
Specific versions change constantly, so it's more useful to think in categories:
| Category | Examples (families) | Strengths | Trade-offs |
|---|---|---|---|
| Frontier, closed | Anthropic Claude, OpenAI GPT, Google Gemini | Highest capability, strong tool use, managed APIs | Per-token cost, data leaves your infrastructure |
| Open-weight | Meta Llama, Mistral, Qwen, DeepSeek, Google Gemma | Self-hostable, fine-tunable, data stays with you | You run the infrastructure; usually a step behind the frontier |
| Reasoning / thinking modes | Extended-thinking variants of most frontier families | Maths, code, multi-step planning, hard analysis | Slower and more output tokens per answer |
| Small and fast | The "mini", "flash" or "haiku" tier of each family | Low latency and cost; great for classification, extraction, routing | Weaker on complex reasoning |
| Specialist | Embedding, reranking, speech, vision and code models | Best at one job (e.g. embeddings for RAG) | Not general-purpose |
Always check the provider's documentation for the current model names, context limits and prices — anything written in a blog post (including this one) will date quickly.
What actually differs between models
- Quality on your task — the only quality measure that matters, and the one no leaderboard can give you.
- Latency — time to first token (how fast it starts) and tokens per second (how fast it finishes). Critical for chat and voice.
- Cost — priced per million input and output tokens; output is usually several times more expensive. Prompt caching and batch APIs can cut costs substantially.
- Context window — and how well the model actually uses information buried in the middle of a long input.
- Tool use and structured output — how reliably it calls functions with correct arguments and returns valid JSON. Make-or-break for agents.
- Multimodality — images, PDFs, audio, video in; images or speech out.
- Deployment and data — available regions, data-retention terms, self-hosting options and licence.
Estimating cost: a worked example
Model prices are quoted per million tokens, with input and output priced separately. Here's how to estimate a monthly bill, using illustrative prices of $3 per million input tokens and $15 per million output tokens:
| Item | Calculation | Result |
|---|---|---|
| Conversations per month | — | 10,000 |
| Input tokens | 10,000 × 2,000 tokens (instructions, context, history) | 20M → $60 |
| Output tokens | 10,000 × 300 tokens | 3M → $45 |
| Total | ≈ $105 / month |
Two things usually dominate: how much context you send on every call (trim it, cache it) and whether you use a large model for requests a small one could handle (route them). Reasoning modes also generate many extra "thinking" tokens, so measure their real cost per task before turning them on everywhere.
Benchmarks: what they measure
A benchmark is a fixed public test set with a scoring method. These are the ones you'll see most often on model launch charts:
| Benchmark | What it measures | Watch out for |
|---|---|---|
| MMLU / MMLU-Pro | Broad knowledge across ~57 subjects, multiple choice (Pro: harder, 10 options) | Original MMLU is saturated — top models all score similarly |
| GPQA Diamond | Graduate-level science questions designed to be "Google-proof" | Small set (~200 questions), so scores are noisy |
| Humanity's Last Exam | Very hard expert questions across many fields | Measures the frontier, says little about everyday tasks |
| AIME, MATH | Competition and school-level mathematics | Tests reasoning in maths specifically |
| HumanEval | Writing short Python functions from a docstring | Saturated and widely leaked into training data |
| SWE-bench Verified | Fixing real GitHub issues in real repositories | Scores depend heavily on the agent scaffold around the model |
| LiveCodeBench | Fresh coding problems collected after model training cut-offs | Designed to resist contamination — a good sign |
| τ-bench, BFCL | Tool and function calling; multi-turn agent tasks with simulated users | Closest public proxy for agent reliability |
| ARC-AGI | Abstract pattern puzzles that are easy for people, hard for models | Measures novel reasoning, not knowledge |
| LMArena (Chatbot Arena) | Human preference in blind side-by-side chats, ranked by Elo | Rewards style and length as well as correctness |
| RULER, needle-in-a-haystack | Finding and using information in long contexts | Simple "needle" tests overstate real long-document ability |
How to read a leaderboard without being misled
- Contamination. If test questions leaked into training data, the model is remembering, not reasoning. Prefer benchmarks with fresh or private questions.
- Saturation. When every top model scores 90%+, the differences are noise.
- Different harnesses. Prompts, number of attempts, "thinking" budgets and agent scaffolds vary between reports. A vendor's number and an independent lab's number for the same model often differ.
- pass@1 vs pass@k. "Solved in one try" and "solved in any of five tries" are very different claims.
- Missing columns. Leaderboards rarely show cost and latency — and a model that is 2% better but 5× the price is often the wrong choice.
- Goodhart's law. Once a benchmark becomes a target, models get optimised for it, and it stops measuring what it was designed to measure.
Use benchmarks to build a shortlist. Use your own evals to make the decision.
Evals: your own benchmark
An eval is a repeatable test of your system on your task: a set of inputs, the expected behaviour, and a way of scoring the output automatically. Benchmarks tell you how a model does on someone else's exam; evals tell you how your product does on yours.
Step 1 — Define what "good" means
Write it down in checkable terms. "Helpful answers" isn't checkable. "Answers the question, cites a source, under 120 words, never invents an order number, escalates refund requests over £500" is.
Step 2 — Build a golden dataset
Collect 50–200 realistic examples: real user messages (anonymised) are best. Include the easy common cases, the awkward edge cases, the ambiguous ones, and a few adversarial ones (prompt injection, off-topic requests, missing information). Each example gets an expected answer or a description of what a good answer must contain.
Step 3 — Choose graders
| Grader | Use it for | Example |
|---|---|---|
| Code-based | Anything objectively checkable — fast, cheap, deterministic | Valid JSON, matches schema, contains the order ID, under N words, correct label |
| LLM-as-judge | Qualities that need judgement | "Is every claim supported by the sources?" "Is the tone appropriate?" |
| Human review | Calibrating the judge, high-stakes outputs, spot checks | A weekly review of 30 sampled conversations |
LLM judges are powerful but biased: they favour longer answers, the first option in a pair, and sometimes their own model family. Mitigate this by giving the judge a specific rubric, asking for a short justification before a pass/fail verdict, swapping answer order in pairwise comparisons, and checking the judge's verdicts against human labels on a sample before trusting it.
Step 4 — Run, compare, decide
Run every candidate model (or prompt, or retrieval setting) over the same dataset and record quality, cost and latency side by side. Here is the core of an eval harness — wire complete() to whichever provider SDK you use:
import json, time
def complete(model: str, prompt: str) -> str:
"""Call your LLM provider here and return the text response."""
raise NotImplementedError
JUDGE_PROMPT = """You are grading a customer-support answer.
Question: {question}
Reference answer: {reference}
Candidate answer: {answer}
Does the candidate answer agree with the reference on every fact,
without adding unsupported claims? Explain in one sentence, then
finish with exactly PASS or FAIL on its own line."""
def run_eval(model: str, dataset_path: str, judge_model: str) -> dict:
# one JSON object per line: {"question": ..., "reference": ..., "must_include": ...}
cases = [json.loads(line) for line in open(dataset_path)]
passed, latencies = 0, []
for case in cases:
start = time.perf_counter()
answer = complete(model, case["question"])
latencies.append(time.perf_counter() - start)
# cheap deterministic checks first
if case.get("must_include") and case["must_include"] not in answer:
continue
verdict = complete(judge_model, JUDGE_PROMPT.format(
question=case["question"], reference=case["reference"], answer=answer))
if verdict.strip().splitlines()[-1].strip() == "PASS":
passed += 1
latencies.sort()
return {
"model": model,
"pass_rate": passed / len(cases),
"p50_latency_s": latencies[len(latencies) // 2],
}
Step 5 — Make it a regression test
Run the eval in CI whenever a prompt, model, retrieval setting or tool definition changes, and fail the build if the pass rate drops below your bar. This is what lets you upgrade to next month's model in an afternoon instead of hoping for the best.
Step 6 — Keep evaluating in production
Offline evals catch regressions; production tells you what you didn't think to test. Log inputs and outputs (respecting privacy), sample a slice for automatic and human grading, collect thumbs-up/down feedback, and add every real failure to the golden dataset. The dataset should grow every week.
Measuring hallucinations
A hallucination is a confident statement that isn't supported by the facts. You can't eliminate them entirely, but you can measure and reduce them:
- Grounded tasks (RAG, summarisation): score faithfulness — the share of claims in the answer supported by the provided sources — with an LLM judge, and spot-check by hand.
- Factual tasks: compare against reference answers and track the rate of unsupported or wrong facts.
- Refusal behaviour: include questions that can't be answered from the data and check the system says so instead of inventing an answer.
- Reduce it with better context, explicit "only from sources" instructions, citations, lower temperature for factual tasks, and structured outputs that leave less room for improvisation.
Evals for RAG and agents
- RAG systems: evaluate retrieval (did we fetch the right chunks? recall@k, MRR) separately from generation (faithfulness, answer relevance). See our RAG architecture guide.
- Agents: score task completion, whether the right tools were chosen with the right arguments, number of steps, recovery from tool errors, and cost per completed task. Check the final state (was the ticket actually created?) rather than just the final message.
- Safety: test prompt injection hidden in documents or emails, attempts to extract other users' data, and requests the system should refuse.
Tools that help
You can start with a script like the one above and a spreadsheet. When you outgrow that, open-source and hosted tools such as promptfoo, DeepEval, Ragas, OpenAI Evals, Langfuse, LangSmith, Braintrust and Arize Phoenix add dataset management, judges, tracing and dashboards. The tool matters far less than having a golden set and running it consistently.
Prompting, RAG or fine-tuning?
Before switching models, check whether the problem is the model at all:
| Problem | Try first | Why |
|---|---|---|
| Wrong format, tone or structure | Better prompt, examples, structured output | Cheapest and fastest to iterate |
| Doesn't know your facts or documents | RAG | Adds knowledge without retraining; supports citations |
| Can't use your systems | Tools / MCP | Gives the model live data and actions |
| Consistent narrow skill at high volume | Fine-tuning a smaller model | Can match a larger model on one task at lower cost |
| Genuinely hard reasoning | A stronger or reasoning model | Some tasks need more capability |
A practical model-selection recipe
- Shortlist three models from benchmarks relevant to your task — say, one frontier, one fast/cheap, one open-weight.
- Run your eval on all three with the same prompt.
- Pick the cheapest, fastest model that clears your quality bar — not the one at the top of the leaderboard.
- Route where it pays. Send simple requests to a small model and escalate hard ones to a larger or reasoning model. Many teams cut costs dramatically this way with no loss in quality.
- Re-run quarterly, or whenever a new model ships. With evals in place, switching is an afternoon's work.
LLM glossary
- Token — the unit models read and write; roughly ¾ of an English word.
- Context window — the maximum tokens a model can consider in one request.
- Temperature — how random the output is; lower for factual tasks, higher for creative ones.
- System prompt — standing instructions that shape the model's behaviour for a conversation.
- Time to first token (TTFT) — how long before the response starts streaming.
- pass@k — the share of problems solved in at least one of k attempts.
- Contamination — test data leaking into training data, inflating benchmark scores.
- Golden dataset — your curated set of inputs and expected outputs used for evals.
Leaderboards tell you who is good at exams. Evals tell you who is good at your job. Only one of those is worth paying for.
Keep reading: see how evals fit into a retrieval system in RAG Architecture, Explained, and how to give your model safe access to tools in MCP vs API: build and deploy an MCP server.
Frequently asked questions
What is the difference between an LLM benchmark and an eval?
A benchmark is a public, fixed test used to compare models on general skills such as knowledge, maths or coding. An eval is your own repeatable test of your system on your task, with your data and your definition of a good answer. Benchmarks shortlist models; evals choose between them.
Which LLM benchmark matters most?
The one closest to your task. For agents, look at tool-use benchmarks such as τ-bench and BFCL; for coding, SWE-bench Verified and LiveCodeBench; for hard reasoning, GPQA Diamond and Humanity's Last Exam. No single benchmark predicts performance on your specific product.
What is LLM-as-a-judge?
LLM-as-a-judge means using a language model to grade another model's output against a rubric, for qualities that simple code can't check, such as faithfulness or tone. It is fast and scalable but biased, so calibrate it against human labels before trusting it.
How many examples do I need for an LLM eval?
Start with 50–200 realistic examples drawn from real usage, including edge cases and a few adversarial inputs. That is enough to catch most regressions. Grow the set every week by adding real failures from production.
How often should I re-evaluate my model choice?
Re-run your evals whenever a prompt, model, tool or retrieval setting changes, and at least quarterly. New models ship frequently, and with an eval suite in place, testing a new one takes hours instead of weeks.
Are open-weight models good enough for production?
Often, yes — especially for focused tasks such as classification, extraction or RAG over your own documents. They let you self-host and keep data in your infrastructure. Run the same eval against an open-weight and a frontier model and let the results decide.