1 Commits
Author SHA1 Message Date
Max Freedom Pollard ff9978f74c fix(search): stop a repeated URL from faking cross-engine agreement
`WebSearch::search` fuses the engines with Reciprocal Rank Fusion, which
sums one term per ranked list. `merge` added a term for every result
instead, so an engine that listed the same canonical URL twice had both
of its ranks counted:

    1/61 + 1/62 = 0.0325

That is what two independent engines agreeing at rank 1 are worth
(2/61 = 0.0328), produced from a single list. Agreement across engines is
the only ranking signal this federation has, and a repeat inside one
engine forges it.

The repeats come from the repository's own canonicalization, not from
exotic input. `canonicalize` in `search/engine.rs` drops the fragment,
the trailing slash and the `utm_*`, `gclid`, `fbclid` and `mc_*`
parameters, and `redirected_target` unwraps the Bing, DuckDuckGo and
Google redirector links, so rows that are visibly distinct on one result
page collapse onto one URL. Nothing dedupes an engine's own list before
`merge` sees it.

`merge` already knew the rule: the engine label was guarded with
`!existing.engines.contains(&engine)`. Put the score behind the same
guard, so each engine contributes its best rank once. Cross-engine
merging and the longest-snippet rule are untouched.
2026-09-07 05:56:57 -04:00