Same engine, same shopper. Only the copy differs.
The consumer side of the same claim, and the platform's ROI rather than one brand's. One catalogue of 30 products runs twice through one retrieval engine — on the left written the way brands write today, on the right agent-ready. The shopper sees three results either way; the question is whether they are the right three.
Who is shopping
Each profile carries at least one hard constraint a listing either answers or does not — soft preference-only profiles cannot separate the arms.
Benchmark across all shoppers
nDCG@3 against ESCI-style graded labels, cached from a full run of all four profiles.
One profile regresses, and it is reported rather than dropped. On “heavier, wide feet” the raw catalogue already surfaced a correct top three and the rewrite reordered it slightly worse. Two profiles sit at ceiling because their binding constraint is price, which survives even in vague copy — the arms separate exactly where the constraint is one raw copy tends to omit (breathability, terrain, fit). A benchmark you always win is not a benchmark.