References
84 items
Extracted from your reading. Status changes save to your account.
| Item | Status | Actions |
|---|---|---|
|
The item points to an arXiv paper and benchmark dataset for evaluating pragmatic multimodal video understanding, which functions as a reference resource worth bookmarking.
dataset
Source
|
saved |
|
|
The core value is a reusable Arabic hallucination-detection QA corpus with annotations, references, and benchmark-style data for verification research.
dataset
Source
|
saved |
|
|
The core value is SocietyBench itself: a released benchmark with timelines, question banks, ground truth, and scoring code for evaluating social-event forecasting.
dataset
Source
|
saved |
|
|
The core value is a released benchmark and automated evaluation pipeline for agent self-evolution, which can be bookmarked and used independently of the paper’s framing.
dataset
Source
|
saved |
|
|
The item’s core value is a benchmark/reference package of curated GeneBench-Pro case studies, prompts, datasets, and supporting materials for biomedical reasoning evaluation.
dataset
Source
|
saved |
|
|
The core item is WorldExam, a concrete benchmark and accompanying paper for evaluating controllable video/world models.
dataset
Source
|
saved |
|
|
The item points to a benchmark suite and accompanying arXiv paper intended as a reference for evaluating language models on wet-lab synthetic chemistry tasks.
dataset
Source
|
saved |
|
|
The item introduces MedPRESS, a concrete medical LLM safety benchmark/dataset that can be used or cited independently of the paper’s narrative.
dataset
Source
|
saved |
|
|
The item points to SCHEDBench, a concrete benchmark/dataset for evaluating LLM constraint faithfulness in scheduling tasks.
dataset
Source
|
saved |
|
|
The item introduces and documents ARB, a benchmark dataset for evaluating AI-text detector robustness, which is a concrete reference resource independent of the paper’s summary findings.
dataset
Source
|
saved |
|
|
The item introduces DungeonBench as a concrete AI evaluation benchmark with documented tracks, rules coverage, and results that researchers can reference or use.
dataset
Source
|
saved |
|
|
The item points to FriendBench, a released benchmark dataset and evaluation reference for dyadic familiarity inference in humans and multimodal models.
dataset
Source
|
saved |
|
|
ExtractBench is a concrete benchmark dataset and evaluation reference for schema-guided enterprise document extraction, with dataset and code links provided.
dataset
Source
|
saved |
|
|
The item primarily points to ACE-Data-0, a concrete multisensory dataset and benchmark for embodied AI research that can be used or bookmarked independently of the paper summary.
dataset
Source
|
saved |
|
|
The core value is the released OSReward benchmark, OS-Shepherd-100K dataset, and related evaluation resources for CUA reward-model research.
dataset
Source
|
saved |
|
|
The item points to an arXiv paper introducing ClinMM-Bench, a concrete multimodal clinical diagnostic benchmark and evaluation reference for researchers.
dataset
Source
|
saved |
|
|
The item primarily points to Desktop-Delta Bench, a concrete benchmark/dataset for evaluating desktop GUI transition understanding in computer-use agents.
dataset
Source
|
saved |
|
|
The core value is a publicly released robotics benchmark, labeled dataset, and leaderboard that researchers can bookmark and use for evaluation.
dataset
Source
|
saved |
|
|
The core value is a publicly released benchmark dataset, evaluation toolkit, and code for assessing VLM understanding of ER/EER diagrams.
dataset
Source
|
saved |
|
|
The core value is a concrete benchmark and dataset for robotic failure analysis, with code/data and evaluation protocols that can be reused or bookmarked.
dataset
Source
|
saved |
|
|
The core value is GEMCo, a releasable proxy dataset and validation reference for privacy-restricted German e-mail counselling research.
dataset
Source
|
saved |
|
|
The item points to E-Bench, a concrete benchmark and accompanying paper for evaluating multi-step tool-use agents, which is best treated as a reference resource.
dataset
Source
|
saved |
|
|
The primary value is the IR275K infrared multi-frame super-resolution benchmark dataset and reproducible evaluation protocol, which are concrete reference resources for researchers.
dataset
Source
|
saved |
|
|
The item points to SceneActBench, a concrete 3D agent-evaluation benchmark and accompanying arXiv paper with tasks, metrics, and results worth bookmarking.
dataset
Source
|
saved |
|
|
The item presents DBA-Bench as a concrete benchmark and accompanying paper for evaluating LLM-based database operations agents.
dataset
Source
|
saved |
|
|
The main value is a released real-world RGB-D biomedical object dataset with annotations and benchmarks that researchers can bookmark and use.
dataset
Source
|
saved |
|
|
The core value is a publicly available Sri Lankan value-alignment dataset, instruction corpus, benchmark, and codebase that others can reuse or benchmark against.
dataset
Source
|
saved |
|
|
The core value is a new benchmark and associated dataset/paper for evaluating expert-level knowledge-intensive visual reasoning.
dataset
Source
|
saved |
|
|
The core value is VEHBench itself, a benchmark/dataset and evaluation reference for LLM-assisted vibration energy harvester design.
dataset
Source
|
saved |
|
|
The core value is a publicly released Bengali YouTube clickbait dataset with benchmark results that researchers can bookmark and use as a reference.
dataset
Source
|
saved |
|
|
The core value is a publicly released Kyrgyz LLM benchmark suite with datasets, evaluation code, and per-model results worth bookmarking for future NLP evaluation work.
dataset
Source
|
saved |
|
|
The core value is a new UAV-DualCog benchmark/dataset and accompanying materials for evaluating multimodal UAV spatio-temporal reasoning.
dataset
Source
|
saved |
|
|
The item presents TikStance as a concrete multimodal dataset for stance-analysis research that readers can reference or use independently.
dataset
Source
|
saved |
|
|
The item points to an open-source medical AI safety benchmark, taxonomy, rubric, and leaderboard that can be used or bookmarked independently of the article’s framing.
dataset
Source
|
saved |
|
|
The item introduces MM-IssueLoc as a concrete benchmark dataset and evaluation protocol for multimodal repository-level issue localization that others can use or cite.
dataset
Source
|
saved |
|
|
The item points to AdvancedMathBench, a concrete benchmark suite with datasets and evaluation resources for advanced mathematical proof generation and verification.
dataset
Source
|
saved |
|
|
The core value is a publicly available benchmark and evaluation framework for visually grounded tool-calling agents, along with reported evaluation results.
dataset
Source
|
saved |
|
|
The item’s main value is a public benchmark dataset and fault-injection framework for DNN training fault detection and diagnosis.
dataset
Source
|
saved |
|
|
The item points to DexVerse as a concrete benchmark suite with tasks, embodiments, demonstrations, and evaluation results that researchers can use or bookmark.
dataset
Source
|
saved |
|
|
The primary value is a publicly available benchmark dataset and associated evaluation reference for long-term multi-object tracking.
dataset
Source
|
saved |
|
|
The core item is a publicly released benchmark and codebase for evaluating proactive agents, which is a concrete reference resource for researchers.
dataset
Source
|
saved |
|
|
The item’s core value is a concrete multitrack transcription dataset and benchmark for evaluating music transcription models.
dataset
Source
|
saved |
|
|
The core value is a concrete benchmark dataset and evaluation reference for incident-centric dashcam VQA in autonomous driving.
dataset
Source
|
saved |
|
|
The item points to a concrete scientific benchmark/dataset and evaluation suite for lineage reasoning and idea generation that researchers could bookmark or use.
dataset
Source
|
saved |
|
|
The core value is an open benchmark release with datasets, transcripts, scorer, and baseline results that others can use to evaluate Chinese news TTS systems.
dataset
Source
|
saved |
|
|
The core value is a released benchmark dataset with annotation guidelines and checkpoints that others can directly use and cite for Marathi NLP work.
dataset
Source
|
saved |
|
|
The core value is a reusable benchmark dataset with accompanying code and baseline results that others can use to evaluate models, making it a concrete reference resource rather than primarily an argument or tutorial.
dataset
Source
|
saved |
|
|
The core value is a concrete benchmark dataset and accompanying evaluation setup that others can use and bookmark for research, despite the paper also containing analysis of model failures.
dataset
Source
|
saved |
|
|
The core value is a concrete dataset resource for training and evaluating counterspeech and RAG systems, not the paper’s narrative alone.
dataset
Source
|
saved |
|
|
The core deliverable is a publicly released multimodal SAR–optical–text dataset with fixed splits and code that others can use as a benchmark and data resource.
dataset
Source
|
saved |
|
|
The core value is a benchmark dataset/evaluation framework (X+Slides) that others can use to assess audience-conditioned slide generation systems.
dataset
Source
|
saved |
|
|
The core value is a released machine-readable corpus of U.S. local ordinances with metadata and related models that researchers can directly use and bookmark as a dataset resource.
dataset
Source
|
saved |
|
|
The core value is a concrete dataset and paper that researchers can use or bookmark for world-modeling work, not primarily the abstract’s narrative itself.
dataset
Source
|
saved |
|
|
The core deliverable is a concrete benchmark dataset and evaluation framework, along with metrics and baselines, that others can use to test workflow-prediction systems.
dataset
Source
|
saved |
|
|
The core value is a benchmark dataset/evaluation framework for visual navigation that others can use to test and compare models, rather than the writeup’s narrative alone.
dataset
Source
|
saved |
|
|
The core value is a concrete dataset with annotations that researchers can use for benchmarking, modeling, and evaluation, rather than the paper’s narrative itself.
dataset
Source
|
saved |
|
|
The core deliverable is a concrete dataset and benchmark (OmniVideo-100K and OmniVideo-Test) that others can use for training and evaluation, rather than the paper’s narrative alone.
dataset
Source
|
saved |
|
|
The core value is a publicly available benchmark dataset and accompanying evaluation setup that others can use to assess and study medical MLLM hallucinations.
dataset
Source
|
saved |
|
|
The core deliverable is a benchmark dataset and evaluation protocol that others can use to compare coding-agent harnesses, making it primarily a reference resource rather than just an analysis article.
dataset
Source
|
saved |
|
|
The core value is a concrete benchmark resource—a physically accurate articulated-object dataset and simulator—that others can use for research and evaluation.
dataset
Source
|
saved |
|
|
The core deliverable is a new benchmark paper and benchmark artifact to evaluate multimodal conversational memory systems, which people would primarily use as a reference resource rather than consume for its narrative alone.
dataset
Source
|
saved |
|
|
The paper presents a real-world dataset aimed at enabling AI-native mobility in 6G by replacing common simulation-based data with measurements from a commercial network.
dataset
Source
|
saved |
|
|
RoboMemArena is a large-scale benchmark addressing robotic memory shortcomings by offering 26 long-horizon tasks (avg >1,000 steps) with multimodal memory annotations and paired real-world evaluations; a VLM-based pipeline composes subtasks and generates trajectories via atomic functions.
dataset
Source
|
saved |
|
|
CADBench is a unified multimodal benchmark for recovering editable CAD programs from 2D/3D inputs, assembled from 18,000 samples across six families and five input modalities with six evaluation metrics.
dataset
Source
|
saved |
|
|
TAVIS is a benchmark for egocentric active vision and anticipatory gaze in imitation learning that provides two complementary suites (Head: 5 tasks; Hands: 3 tasks) on GR1T2 and Reachy2 torsos in IsaacLab, plus ~2,200 demonstrations.
dataset
Source
|
saved |
|
|
OpenAI Five Benchmark on August 5, 2018 staged a live best-of-three match (stream 12:30pm PT, main event 1:15pm PT) in San Francisco against 99.95th-percentile Dota players.
dataset
Source
|
saved |
|
|
SimpleQA is a compact factuality benchmark released 2024-10-30 consisting of 4,326 short, single-answer questions whose reference answers were independently produced by two AI trainers and spot-checked by a third (94.4% agreement; ~3% error).
dataset
Source
|
saved |
|
|
Benchmark U.S. Construction Pay
The Michael Page 2026 U.S.
dataset
Source
|
saved |
|
|
COOPER women's NCAA basketball ratings
COOPER is Silver Bulletin’s new Elo‑based power rating for women’s Division I basketball, released March 12, 2026 with ratings through March 11.
dataset
Source
|
saved |
|
|
The French-YMCA corpus is a new linguistic resource targeting children and adolescents, comprising 39,200 texts and 22,471,898 words drawn from diverse sources with consistent grammar and spelling.
dataset
Source
|
saved |
|
|
BrowseComp is a 1,266-item OpenAI benchmark (Apr 10, 2025) of intentionally obscure, single-answer browsing tasks created by human trainers with strict checks to ensure difficulty and verifiability.
dataset
Source
|
saved |
|
|
Humanity’s Last Exam is a 2,500-question test created by scientists to probe deep, cross-disciplinary expert knowledge across mathematics, humanities, natural sciences, ancient languages and niche subfields.
dataset
Source
|
saved |
|
|
IndQA is a new OpenAI benchmark (released 2025-11-03) of 2,278 culturally grounded, reasoning-focused questions in 12 Indian languages across 10 domains, authored and reviewed by 261 experts.
dataset
Source
|
saved |
|
|
Qwen 3.5 benchmark comparison website (published 2026-02-25) aggregates verified benchmark scores and head-to-head infographics for models including GPT-5.2, Claude 4.5 Opus, Gemini-3 Pro, Qwen3-Max-Thinking, K2.5-1T-A32B, GPT-OSS-120B and Qwen3.5 variants (397B, 235B, 122B, 35B, 27B).
dataset
Source
|
saved |
|
|
Epoch AI’s new explorer is a useful primary-source-style attempt to quantify the global installed base of AI accelerators, an increasingly important constraint for model training and deployment.
dataset
Source
|
saved |
|
|
MTRAG-UN is a 666-task benchmark (released Feb 2026 by Rosenthal et al., IBM) for multi-turn retrieval-augmented generation that emphasizes hard conversational phenomena: unanswerable, underspecified, non‑standalone questions and unclear responses.
dataset
Source
|
saved |
|
|
SPARTA is a benchmark and prompt-template suite aimed at scalable, principled evaluation of tree-structured multi-hop question answering that must operate over both natural-language text and structured tables.
dataset
Source
|
saved |
|
|
SumTablets is a curated, machine-readable collection that pairs Unicode cuneiform glyph sequences with scholarly Latin‑alphabet transliterations for 91,606 Sumerian tablets (6.97M glyphs).
dataset
Source
|
saved |
|
|
PleIAs/SYNTH is a synthetic Q&A/reasoning dataset posted on Hugging Face for text-generation tasks, roughly 3.8k examples and associated with a 0.3B tag and frequent use of the qwen-3-8b-memorization label.
dataset
Source
|
saved |
|
|
Common Corpus is a two-trillion-token open dataset (announced 2025-06-02) aggregating uncopyrighted or permissively-licensed text and code for LLM pre-training.
dataset
Source
|
saved |
|
|
Common Corpus is a 2 trillion-token, truly open multilingual dataset released 2026-01-26 (OpenReview id 0wSlFpMsGb) assembled from uncopyrighted or permissively licensed material.
dataset
Source
|
saved |
|
|
HealthBench is a new OpenAI benchmark (released 2025-05-12) of 5,000 realistic health dialogues with physician-crafted rubrics (48,562 criteria) created by 262 physicians across 60 countries.
dataset
Source
|
saved |
|
|
PVIR introduces a 95-video benchmark to probe physics-aware video instance removal, focusing on causal side effects (reflections, illumination, shadows) that prior benchmarks ignore.
dataset
Source
|
saved |
|
|
Free to Use Datasets and Tools to Benchmark Rates and Analyze Affordability was presented as a NARUC CPI innovation webinar (published 2026-03-30) and moderated by Commissioner Christine Needo (appointed February 3, 2024).
dataset
Source
|
saved |
|