Selected work / ToolRet

05 / INDEPENDENT RETRIEVAL RESEARCH

ToolRet logo

ToolRet

When does reranking earn its cost? An empirical study of conditional tool reranking across a 37,292-tool catalog, with untouched confirmation queries and explicit quality tolerances.

TOOLRET / SYSTEM OVERVIEW

BM25 + denseRRFRouterRerankConditional work, measured quality
PythonPyTorchMiniLM

The problem

Reranking can improve retrieval, but always running a cross-encoder adds cost. The research question is whether a learned router can skip enough calls while staying within a quality tolerance fixed before confirmation.

Engineering decisions

  • Combined BM25 and MiniLM dense retrieval with reciprocal-rank fusion over 37,292 tools.
  • Built a seven-feature ridge router for conditional cross-encoder reranking.
  • Separated development, calibration, and confirmation by relevant-tool components and froze parameters before confirmation.
  • Compared 20-seed global and source-matched random controls and retained a repository-hosted study report and independent audit.

What the evidence shows

On 1,500 untouched confirmation queries, the router used 23.4% fewer cross-encoder calls while meeting a predeclared 0.01 nDCG@10 loss tolerance. Quality superiority was not established. This is independent undergraduate empirical research, not a peer-reviewed publication.

The study report, confirmation results, and independent audit are retained in the public repository.