Evaluating search relevance without human raters is a persistent challenge for search teams. One developer documented an experiment using LLMs as pairwise judges, testing how different prompting strategies affect agreement with human labels. The work began with a simple question: can an LLM reliably say which of two products is more relevant to a given query?

The starting point was a forced decision. Given a query "leather chairs" and two product names, the LLM was asked to pick LHS or RHS. Over 1,000 pairs, this approach yielded 75.08% precision with 100% recall. Every pair was labeled, but a significant fraction disagreed with human raters.

A second variant allowed the model to refuse. If evidence was insufficient, the LLM could output "Neither." This improved precision to 85.38%, but recall crashed to 17.10%. Only a subset of pairs received labels, and of those, 85.38% matched human judgments.

An important sanity check came from double‑checking each decision. The same query and products were presented twice, swapping LHS and RHS positions. If the LLM was consistent, the first result was kept; inconsistent pairs were marked "Neither." This alone raised precision to 87.99% while maintaining 65.80% recall.

Combining the "allow neither" option with the double‑check produced the highest precision—90.76%—but recall fell to 11.90%. The tradeoff is clear: teams willing to accept more labels tolerate more errors, while teams that can afford to miss cases can achieve very high accuracy by requiring confidence and consistency.

The author also tested other product fields as relevance signals. Using only the product class (e.g., "Beds" vs. "Kids Beds") gave results comparable to the name‑only approach. The full categorization hierarchy (e.g., "Outdoor furniture > Seating > Adirondak Chairs") slightly improved recall over class alone but still lagged behind product names. Product descriptions were the noisiest field; even when forced to decide, the LLM sometimes declined unprompted, and precision remained the lowest across all categories.

ApproachDont check--check-both-ways
Force (name)75.08% / 100%87.99% / 65%
Allow Neither (name)85.38% / 17.10%90.76% / 11.90%
Force (class)70.5% / 100%87.76% / 58.0%
Allow Neither (class)87.01% / 17.70%84.47% / 10.3%
Force (hierarchy)74.6% / 100%86.1% / 69.70%
Allow Neither (hierarchy)85.71% / 18.20%89.91% / 10.8%
Force (description)70.31% / 98.70%76.58% / 72.60%
Allow Neither (description)79.21% / 10.10%83.02% / 5.3%

The author concludes that the choice of approach depends on the use case. If coverage matters more than accuracy, the forced decision without double‑checking is sufficient. If errors are costly and some missed cases are acceptable, allowing "Neither" and checking both ways produces the most reliable results.

Beyond the experiment, the author is launching a course titled "Cheat at Search with Agents" in October. The course will cover using agents in search, building better RAG systems, and applying LLMs for query understanding. Enrollment is open at the Maven platform.