0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilor
dفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy
jmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈
qtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0
x▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بd
◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgj
∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknq
فehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux
ilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇
psvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞
wz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0اف
◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفi
π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmp
cلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtw
hknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆
orux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π
vy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xc
░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افeh
∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilo
بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsv
gjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░
nqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑
ux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1ب
▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلgjmpsvy▒◇∆π1بdفilorux▓◆∇∑0افehknqtwz░◈Ξ∞xcلg
rehan@neural-mesh :~
$
R
Rehan Rao
all insights
AI · Evaluation 2025 · 07 12 min

LLM evaluation matrices that actually catch regressions.

Vibes-based eval kills production models. Here's the multi-axis matrix — reasoning, adversarial, hallucination, domain fit — that scaled a 30% quality improvement.

Rehan Rao
AI & Backend Systems Engineer
architectureeval.matrix
ReasoningAdversarialHallucinationDomain FitLatencygpt-4o9278848870claude-3.59082888474llama-3.18070747690

“It feels better” is not an evaluation. It is a story you tell yourself before a customer finds a regression. The most reliable way I have found to ship LLM upgrades safely is a five-axis matrix with hard numerical gates.

Why vibes-eval fails

Manual spot-checks are biased by recency and novelty. A new model that phrases things more confidently reads as smarter — even when its factual accuracy dropped. You need a harness that answers a single question: is this checkpoint safe to promote?

The five axes

  • Reasoning — multi-step arithmetic, deduction, tool selection.
  • Adversarial — prompt injections, jailbreaks, misleading context.
  • Hallucination — closed-book questions with a known answer set.
  • Domain fit — real user queries from production traces, hand-labeled.
  • Latency — P50/P95 wall-clock at target concurrency.

The harness

Every axis is a dataset of (input, expected, grader) triples. Graders are either deterministic (regex, JSON schema, exact match) or an LLM judge with a rubric. Never mix the two on the same axis — you cannot compare scores across grading methods.

@dataclass
class Case:
    axis: str
    prompt: str
    expected: Any
    grader: Callable[[str, Any], float]  # -> 0..1

def run(model, cases):
    return {
        axis: mean(c.grader(model(c.prompt), c.expected)
                   for c in cases if c.axis == axis)
        for axis in AXES
    }

Regression gates

A candidate model must clear every axis by ≥ the baseline − 2%. A gain on reasoning does not excuse a hallucination regression. Ship the matrix in CI; block the deploy if any axis regresses.

caveat
Do not average axes into a single score. Averages hide the exact failure mode you care about — a model that hallucinates 15% more but reasons 20% better still looks “ better” on the mean.

Results

On the last rollout I ran, this matrix caught two regressions that human review missed — both hallucination spikes on domain-specific numerical claims. The final promoted model delivered ~30% aggregate quality lift while holding latency flat.