A commit message in my hybrid-search implementation claimed that unequal weights made exact ties impossible. The scoring formula provides a counterexample, and the code already had a tie-break for it.
Unequal weights mean an exact tie cannot occur, so no positional tie-break decides which retriever wins a disagreement.
Checking the possible scores exposed the mistake. Giving two search methods different weights changes their influence, but doesn’t guarantee that every document gets a different score.
Combining two kinds of search
Vector search alone can miss exact identifiers and error strings. A docs query like TURSO_DB_URL benefits from keyword matching, while a question phrased in everyday language can benefit from vector similarity. The search endpoint runs both and combines their ranked results.
The method is weighted reciprocal rank fusion, or RRF. It uses each document’s position in the result lists rather than comparing the keyword and vector scores directly. Those scores use different scales, so working with ranks avoids having to calibrate them against each other.
Each list contributes weight / (k + rank) to a document’s combined score, with ranks starting at 1. This configuration gives vector results slightly more weight:
export const RRF_K = 60;
export const FUSION_WEIGHTS = {
keyword: 0.85,
vector: 1,
} as const;
Different weights prevent two documents found only at the same rank in different lists from receiving the same contribution. That doesn’t cover documents at different ranks, or documents found by both search methods.
The value k=60 comes from experiments in the original RRF paper. It controls how much rank differences affect the score. It’s a reasonable starting point to evaluate for this project, rather than a guarantee about ranking quality or ties.
A counterexample
With ten requested results, each search method supplies up to thirty candidates. Two positions inside that candidate pool give a simple mathematical tie:
- A keyword-only hit at rank 8 scores
0.85 / 68, or 0.0125. - A vector-only hit at rank 20 scores
1 / 80, also 0.0125.
Treating 0.85 as the exact fraction 17/20, both contributions reduce to 1/80. Unequal weights haven’t prevented equality. These are candidate positions, though, not a promise that either document appears in the final ten results.
There is a wrinkle when checking the implementation: JavaScript uses floating-point numbers. Those two divisions round slightly differently, so they don’t compare equal with ===. The mathematical counterexample alone doesn’t demonstrate that the code’s tie-break runs.
Another pair does. Suppose document A is third in the vector results and twenty-fourth in the keyword results. Document B is twentieth in the vector results and third in the keyword results:
const scoreA = 1 / (60 + 3) + 0.85 / (60 + 24);
const scoreB = 1 / (60 + 20) + 0.85 / (60 + 3);
scoreA === scoreB; // true
Both evaluate to 0.02599206349206349. This pair ties mathematically and in the JavaScript arithmetic used by the implementation.
Keeping the tie-break
The sort already compares document IDs when scores are equal. That gives tied results a defined order, independent of which search list added them first.
The incorrect explanation didn’t cause an ordering bug, but it gave a bad reason to consider that safeguard unnecessary. Removing the ID comparison would make equal-score ordering depend on the order supplied to the sort. It wouldn’t automatically change on every request, but it would leave that decision to how the result lists were assembled.
This is also a useful regression-test case: give two documents those ranks, confirm their scores match, and check that their IDs decide the order. It tests a specific behavior that the weight configuration doesn’t guarantee.
I’m keeping the document-ID tie-break. Different weights change the ranking, but they don’t remove the need to decide what happens when scores match.
Sources
- Reciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods (PDF) — Cormack, Clarke, Büttcher, SIGIR 2009, the ranking formula and experiments behind k=60
- The commit — FTS5 keyword results fused with vector search
- Fusion implementation — score calculation and the document-ID tie-break
I’d appreciate a follow. You can subscribe with your email below. The emails go out once a week, or you can find me on Mastodon at @[email protected].