What Does It Mean When a Table Has Dashes for a Model?

I’ve spent the better part of a dozen years watching teams scramble to integrate the "latest and greatest" model, only to realize three months later that their production environment is hemorrhaging errors. One of the most common red flags I see—which teams almost always ignore—is the infamous "dash" (—) in a comparison table. You know the one: a row for OpenAI’s latest, a column for Anthropic’s, and a glaring, empty dash where a metric should be.

When you see no published benchmark data for a specific model-task pairing, your first instinct should be skepticism, not assumption. Are they hiding something? Did they just not run the test? Did the model refuse to answer the prompt entirely?

The Anatomy of the Missing Metric

In the world of AI evaluation, a dash is never neutral. It is an data point in its own right. When a model lacks a score on a leaderboard—like the Vectara HHEM Leaderboard or the Artificial Analysis AA-Omniscience suite—it usually boils down to one of three failures:

    Refusal Behavior: The model identified the prompt as "out of scope" or "unsafe" and returned a canned response, resulting in a null score rather than a wrong one. Incompatibility: The benchmarking framework relies on a specific output format (like JSON) that the model struggles to adhere to, leading to a parsing crash. Strategic Omission: The provider chose not to submit the model to that specific test because the performance profile is known to be sub-optimal for that specific task.

Before you trust any leaderboard, you have to ask: What exactly was measured? If Google’s latest iteration doesn't have a score for citation accuracy, it isn't a "lack of information"—it’s a boundary condition you need to account for.

The "Hallucination" Trap

Let’s be blunt: Hallucinations are currently unavoidable in generative AI. If anyone tells you they have "near zero hallucinations," they are either selling you a pipe dream or measuring the wrong thing. The goal isn't to eliminate them; it's to quantify the risk so you can build guardrails.

Metric Type What It Actually Measures Common Failure Mode Summarization Faithfulness Does the output follow the source text? Ignoring contradictory information. Knowledge Reliability Is the latent knowledge accurate? Mixing training data with prompt data. Citation Accuracy Do the sources support the claims? Hallucinating document IDs.

Why Benchmark Mismatch is a Silent Killer

Teams love to cherry-pick a score from one benchmark and apply it to their specific use case. This is how projects go off the rails. A model might perform beautifully on general reasoning benchmarks but fail catastrophically at "groundedness"—the ability to stay strictly within the context of a provided document.

image

When comparing Artificial Analysis AA-Omniscience results against internal testing, look for the delta. If a model ranks high on broad, chat-based benchmarks but has dashes on narrow, retrieval-augmented generation (RAG) benchmarks, it tells you exactly where that model’s "truth-seeking" mechanisms break down.

The Danger of Assuming "No Data" Means "High Performance"

The most dangerous assumption in enterprise AI is thinking, "Well, the base model is smart, so it should handle RAG just fine." This is the primary reason developers get burned. When you see a missing metric, you are looking at a blind spot.

image

Refusal behavior vs. wrong-answer behavior is the nuance most teams ignore. If a model hallucinates a fact, that’s a data error. If a model refuses to answer because it isn't "confident" enough, that’s a system behavior. A high hallucination rate is often just a symptom of a model that doesn't know when to say, "I don't know."

Checklist for Interpreting Missing Data

Verify the Benchmarking Tool: Is it evaluating the model's chat capability or its ability to act as a logic engine? Look for Refusal Rates: If a model has high refusal rates, its "accuracy" score will look artificially high because it simply skips the hard questions. Cross-Reference: Don't rely on a single vendor's public leaderboard. Use multi-source verification. Test for "I Don't Know": If you are building a product where truth matters, test if the model is capable of saying "I don't know" before you trust its performance on a "knowledge-heavy" test. https://reliabless.com/ai-that-works-like-having-five-experts-review-your-decision-simultaneously/

Final Thoughts: Don't Settle for One Score

No single leaderboard settles the debate. Whether it's OpenAI, Anthropic, or Google, every provider is optimizing for different RLHF (Reinforcement Learning from Human Feedback) objectives. Some optimize for helpfulness (which encourages hallucinations), while others optimize for safety (which encourages high refusal rates).

When you see those dashes, don't just fill them in with your own optimism. Treat them as a mandate to run your own custom evaluation. If the data isn't there, the responsibility for finding it is yours. In enterprise AI, what you don't measure is exactly what will eventually break your product.