Shopper mode30-SKU catalogue · two-stage retrieval

Same engine, same shopper. Only the copy differs.

The consumer side of the same claim, and the platform's ROI rather than one brand's. One catalogue of 30 products runs twice through one retrieval engine — on the left written the way brands write today, on the right agent-ready. The shopper sees three results either way; the question is whether they are the right three.

1
Embed
shortlist 8 of 30
2
Rerank
LLM picks the top 3
3
Grade
nDCG@3 vs ESCI labels
1

Who is shopping

Each profile carries at least one hard constraint a listing either answers or does not — soft preference-only profiles cannot separate the arms.

3

Benchmark across all shoppers

nDCG@3 against ESCI-style graded labels, cached from a full run of all four profiles.

4 profiles · 30 SKUs
Budget beginner
Raw
1.000
Agent-ready
1.000
ceiling
Humid half-marathon
Raw
0.170
Agent-ready
0.339
+0.169
Heavier, wide feet
Raw
1.000
Agent-ready
0.922
-0.078
Weekend trail
Raw
0.823
Agent-ready
1.000
+0.177
Mean nDCG@3raw 0.748→agent-ready 0.815+0.067

One profile regresses, and it is reported rather than dropped. On “heavier, wide feet” the raw catalogue already surfaced a correct top three and the rewrite reordered it slightly worse. Two profiles sit at ceiling because their binding constraint is price, which survives even in vague copy — the arms separate exactly where the constraint is one raw copy tends to omit (breathability, terrain, fit). A benchmark you always win is not a benchmark.