Indian Flag
Government Of India
A-
A
A+
ORGANISATION
Long-Context-Retrieval-Risk-Dataset-Open-Weight-LLMs

Long-Context-Retrieval-Risk-Dataset-Open-Weight-LLMs

Simulated lost-in-the-middle retrieval accuracy across needle position, context length, and extension method for 168 open-weight model configs (Gemma, Gemma 2, Llama 3.1, Mistral), grounded in the Lost in the Middle (2023), YaRN, StreamingLLM, and RULER literature (2023-2024). Forward-looking task predicts 4x-extension degradation from an in-distribution native-length benchmark only.

About Dataset

This dataset simulates how retrieval and QA accuracy vary with the position of relevant information inside the context window and with how far context length is pushed past a model's native trained length, for 168 configs spanning seven approximate open-weight model configs, four context-extension methods (none, Position Interpolation, YaRN, and the training-free Self-Extend method), and whether the serving stack preserves StreamingLLM-style attention-sink tokens. Retrieval accuracy combines a symmetric U-shaped position-in-context penalty calibrated to the shape reported in Lost in the Middle (2023), an attention-sink boost, and a method-specific exponential decay in extension ratio calibrated to each cited method's reported relative degradation. All ten grounding sources, spanning 2021 to 2024, and every equation are documented in long_context_theory.py and the README Research basis. Simulated model configs reflect 2024-2025 open-weight releases (Gemma 2, Llama 3.1, Mistral); the position-encoding and context-extension mechanisms grounding the equations trace back to RoPE in 2021.

Purpose of Dataset

Lets Teams Choosing A Context-extension Method Estimate, From A Cheap In-distribution Benchmark Run At Or Below A Model's Native Trained Context Length, Whether The Same Configuration Will Still Retrieve Information Reliably Once Pushed To 4x That Length, Before Running A Full Long-context Evaluation Suite.

Activity Overview Activity Overview

  • Downloads0
  • Downloads 0
  • File Size 4.02 MB
  • Views 4

Tags Tags

  • Large Language Model

License Control License Control

Attribution 4.0 International (CC BY- 4.0)

spec-decoding-dataset.json ( 4.02 MB )


To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.

Data Quality Score BetaData Quality Score Beta

Version Control Version Control

FolderVersion 1(4.02 MB)
  • JAI MALI·2 day(s) ago
    • application/json
      spec-decoding-dataset.json