Gabín, JorgeParapar, JavierWang, Xi2026-08-252026-08-252026Jorge Gabín, Javier Parapar, and Xi Wang. 2026. Beyond Top-𝑒: SimulationBased Interactive Evaluation for Query Suggestions. In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’26), July 20–24, 2026, Melbourne, VIC, Australia. ACM, New York, NY, USA, pp. 3167 - 3174. https://doi.org/10.1145/3805712.3808618979-8-4007-2599-9https://hdl.handle.net/2183/49086Presented at: SIGIR '26: Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, July 20–24, 2026, Melbourne, Australia[Abstract]: Evaluating query suggestion systems in a manner that reflects real-world query formulation remains a persistent challenge. Most offline methodologies adopt static assumptions, such as users accepting all or the top-e suggestions, ignoring the inherently selective and intent-driven nature of interactive search. While online experiments provide realistic behavioural signals, they are costly, difficult to scale, and often irreproducible. To bridge this gap, we introduce SIQSE (Simulation-based Interactive Query Suggestion Evaluation), a framework that models query reformulation as an interactive selection task performed by a simulated user. In SIQSE, a Large Language Model (LLM) acts as a surrogate user that progressively selects suggestions according to contextual relevance and explicit search intent. Unlike static offline protocols, this simulation captures the iterative and selective dynamics of real query formulation. Our contributions are twofold. First, we develop and validate an LLM-based selection model, systematically analysing how varying levels of intent information and selection strategies affect its ability to approximate human selection behaviour. Second, we employ this selector to benchmark multiple query suggestion systems across diverse datasets under interactive conditions. Importantly, while the selector is LLM-based, the final evaluation is computed exclusively through ranking-based effectiveness metrics over the rankings produced by selected expansions, ensuring that system performance reflects retrieval quality rather than alignment with the surrogate user model. By modelling round-based interaction while maintaining metric independence, SIQSE offers a scalable, reproducible evaluation paradigm that brings offline assessment closer to the complexity of real-world search behaviour. To facilitate adoption and reproducibility, we release SIQSE as an open-source Python library.engAttribution 4.0 Internationalhttp://creativecommons.org/licenses/by/4.0/Interactive Information RetrievalInteractive Query Suggestion EvaluationSimulation-based EvaluationOffline EvaluationBeyond Top-e: Simulation-Based Interactive Evaluation for Query Suggestionsconference outputopen access10.1145/3805712.3808618