Lists · references

References

84 items

Extracted from your reading. Status changes save to your account.

Item Status Actions
The item points to an arXiv paper and benchmark dataset for evaluating pragmatic multimodal video understanding, which functions as a reference resource worth bookmarking.
dataset Source
saved
The core value is a reusable Arabic hallucination-detection QA corpus with annotations, references, and benchmark-style data for verification research.
dataset Source
saved
The core value is SocietyBench itself: a released benchmark with timelines, question banks, ground truth, and scoring code for evaluating social-event forecasting.
dataset Source
saved
The core value is a released benchmark and automated evaluation pipeline for agent self-evolution, which can be bookmarked and used independently of the paper’s framing.
dataset Source
saved
The item’s core value is a benchmark/reference package of curated GeneBench-Pro case studies, prompts, datasets, and supporting materials for biomedical reasoning evaluation.
dataset Source
saved
The core item is WorldExam, a concrete benchmark and accompanying paper for evaluating controllable video/world models.
dataset Source
saved
The item points to a benchmark suite and accompanying arXiv paper intended as a reference for evaluating language models on wet-lab synthetic chemistry tasks.
dataset Source
saved
The item introduces MedPRESS, a concrete medical LLM safety benchmark/dataset that can be used or cited independently of the paper’s narrative.
dataset Source
saved
The item points to SCHEDBench, a concrete benchmark/dataset for evaluating LLM constraint faithfulness in scheduling tasks.
dataset Source
saved
The item introduces and documents ARB, a benchmark dataset for evaluating AI-text detector robustness, which is a concrete reference resource independent of the paper’s summary findings.
dataset Source
saved
The item introduces DungeonBench as a concrete AI evaluation benchmark with documented tracks, rules coverage, and results that researchers can reference or use.
dataset Source
saved
The item points to FriendBench, a released benchmark dataset and evaluation reference for dyadic familiarity inference in humans and multimodal models.
dataset Source
saved
ExtractBench is a concrete benchmark dataset and evaluation reference for schema-guided enterprise document extraction, with dataset and code links provided.
dataset Source
saved
The item primarily points to ACE-Data-0, a concrete multisensory dataset and benchmark for embodied AI research that can be used or bookmarked independently of the paper summary.
dataset Source
saved
The core value is the released OSReward benchmark, OS-Shepherd-100K dataset, and related evaluation resources for CUA reward-model research.
dataset Source
saved
The item points to an arXiv paper introducing ClinMM-Bench, a concrete multimodal clinical diagnostic benchmark and evaluation reference for researchers.
dataset Source
saved
The item primarily points to Desktop-Delta Bench, a concrete benchmark/dataset for evaluating desktop GUI transition understanding in computer-use agents.
dataset Source
saved
The core value is a publicly released robotics benchmark, labeled dataset, and leaderboard that researchers can bookmark and use for evaluation.
dataset Source
saved
The core value is a publicly released benchmark dataset, evaluation toolkit, and code for assessing VLM understanding of ER/EER diagrams.
dataset Source
saved
The core value is a concrete benchmark and dataset for robotic failure analysis, with code/data and evaluation protocols that can be reused or bookmarked.
dataset Source
saved
The core value is GEMCo, a releasable proxy dataset and validation reference for privacy-restricted German e-mail counselling research.
dataset Source
saved
The item points to E-Bench, a concrete benchmark and accompanying paper for evaluating multi-step tool-use agents, which is best treated as a reference resource.
dataset Source
saved
The primary value is the IR275K infrared multi-frame super-resolution benchmark dataset and reproducible evaluation protocol, which are concrete reference resources for researchers.
dataset Source
saved
The item points to SceneActBench, a concrete 3D agent-evaluation benchmark and accompanying arXiv paper with tasks, metrics, and results worth bookmarking.
dataset Source
saved
The item presents DBA-Bench as a concrete benchmark and accompanying paper for evaluating LLM-based database operations agents.
dataset Source
saved
The main value is a released real-world RGB-D biomedical object dataset with annotations and benchmarks that researchers can bookmark and use.
dataset Source
saved
The core value is a publicly available Sri Lankan value-alignment dataset, instruction corpus, benchmark, and codebase that others can reuse or benchmark against.
dataset Source
saved
The core value is a new benchmark and associated dataset/paper for evaluating expert-level knowledge-intensive visual reasoning.
dataset Source
saved
The core value is VEHBench itself, a benchmark/dataset and evaluation reference for LLM-assisted vibration energy harvester design.
dataset Source
saved
The core value is a publicly released Bengali YouTube clickbait dataset with benchmark results that researchers can bookmark and use as a reference.
dataset Source
saved
The core value is a publicly released Kyrgyz LLM benchmark suite with datasets, evaluation code, and per-model results worth bookmarking for future NLP evaluation work.
dataset Source
saved
The core value is a new UAV-DualCog benchmark/dataset and accompanying materials for evaluating multimodal UAV spatio-temporal reasoning.
dataset Source
saved
The item presents TikStance as a concrete multimodal dataset for stance-analysis research that readers can reference or use independently.
dataset Source
saved
The item points to an open-source medical AI safety benchmark, taxonomy, rubric, and leaderboard that can be used or bookmarked independently of the article’s framing.
dataset Source
saved
The item introduces MM-IssueLoc as a concrete benchmark dataset and evaluation protocol for multimodal repository-level issue localization that others can use or cite.
dataset Source
saved
The item points to AdvancedMathBench, a concrete benchmark suite with datasets and evaluation resources for advanced mathematical proof generation and verification.
dataset Source
saved
The core value is a publicly available benchmark and evaluation framework for visually grounded tool-calling agents, along with reported evaluation results.
dataset Source
saved
The item’s main value is a public benchmark dataset and fault-injection framework for DNN training fault detection and diagnosis.
dataset Source
saved
The item points to DexVerse as a concrete benchmark suite with tasks, embodiments, demonstrations, and evaluation results that researchers can use or bookmark.
dataset Source
saved
The primary value is a publicly available benchmark dataset and associated evaluation reference for long-term multi-object tracking.
dataset Source
saved
The core item is a publicly released benchmark and codebase for evaluating proactive agents, which is a concrete reference resource for researchers.
dataset Source
saved
The item’s core value is a concrete multitrack transcription dataset and benchmark for evaluating music transcription models.
dataset Source
saved
The core value is a concrete benchmark dataset and evaluation reference for incident-centric dashcam VQA in autonomous driving.
dataset Source
saved
The item points to a concrete scientific benchmark/dataset and evaluation suite for lineage reasoning and idea generation that researchers could bookmark or use.
dataset Source
saved
The core value is an open benchmark release with datasets, transcripts, scorer, and baseline results that others can use to evaluate Chinese news TTS systems.
dataset Source
saved
The core value is a released benchmark dataset with annotation guidelines and checkpoints that others can directly use and cite for Marathi NLP work.
dataset Source
saved
The core value is a reusable benchmark dataset with accompanying code and baseline results that others can use to evaluate models, making it a concrete reference resource rather than primarily an argument or tutorial.
dataset Source
saved
The core value is a concrete benchmark dataset and accompanying evaluation setup that others can use and bookmark for research, despite the paper also containing analysis of model failures.
dataset Source
saved
The core value is a concrete dataset resource for training and evaluating counterspeech and RAG systems, not the paper’s narrative alone.
dataset Source
saved
The core deliverable is a publicly released multimodal SAR–optical–text dataset with fixed splits and code that others can use as a benchmark and data resource.
dataset Source
saved
The core value is a benchmark dataset/evaluation framework (X+Slides) that others can use to assess audience-conditioned slide generation systems.
dataset Source
saved
The core value is a released machine-readable corpus of U.S. local ordinances with metadata and related models that researchers can directly use and bookmark as a dataset resource.
dataset Source
saved
The core value is a concrete dataset and paper that researchers can use or bookmark for world-modeling work, not primarily the abstract’s narrative itself.
dataset Source
saved
The core deliverable is a concrete benchmark dataset and evaluation framework, along with metrics and baselines, that others can use to test workflow-prediction systems.
dataset Source
saved
The core value is a benchmark dataset/evaluation framework for visual navigation that others can use to test and compare models, rather than the writeup’s narrative alone.
dataset Source
saved
The core value is a concrete dataset with annotations that researchers can use for benchmarking, modeling, and evaluation, rather than the paper’s narrative itself.
dataset Source
saved
The core deliverable is a concrete dataset and benchmark (OmniVideo-100K and OmniVideo-Test) that others can use for training and evaluation, rather than the paper’s narrative alone.
dataset Source
saved
The core value is a publicly available benchmark dataset and accompanying evaluation setup that others can use to assess and study medical MLLM hallucinations.
dataset Source
saved
The core deliverable is a benchmark dataset and evaluation protocol that others can use to compare coding-agent harnesses, making it primarily a reference resource rather than just an analysis article.
dataset Source
saved
The core value is a concrete benchmark resource—a physically accurate articulated-object dataset and simulator—that others can use for research and evaluation.
dataset Source
saved
The core deliverable is a new benchmark paper and benchmark artifact to evaluate multimodal conversational memory systems, which people would primarily use as a reference resource rather than consume for its narrative alone.
dataset Source
saved
The paper presents a real-world dataset aimed at enabling AI-native mobility in 6G by replacing common simulation-based data with measurements from a commercial network.
dataset Source
saved
RoboMemArena is a large-scale benchmark addressing robotic memory shortcomings by offering 26 long-horizon tasks (avg >1,000 steps) with multimodal memory annotations and paired real-world evaluations; a VLM-based pipeline composes subtasks and generates trajectories via atomic functions.
dataset Source
saved
CADBench is a unified multimodal benchmark for recovering editable CAD programs from 2D/3D inputs, assembled from 18,000 samples across six families and five input modalities with six evaluation metrics.
dataset Source
saved
TAVIS is a benchmark for egocentric active vision and anticipatory gaze in imitation learning that provides two complementary suites (Head: 5 tasks; Hands: 3 tasks) on GR1T2 and Reachy2 torsos in IsaacLab, plus ~2,200 demonstrations.
dataset Source
saved
OpenAI Five Benchmark on August 5, 2018 staged a live best-of-three match (stream 12:30pm PT, main event 1:15pm PT) in San Francisco against 99.95th-percentile Dota players.
dataset Source
saved
SimpleQA is a compact factuality benchmark released 2024-10-30 consisting of 4,326 short, single-answer questions whose reference answers were independently produced by two AI trainers and spot-checked by a third (94.4% agreement; ~3% error).
dataset Source
saved
Benchmark U.S. Construction Pay
The Michael Page 2026 U.S.
dataset Source
saved
COOPER women's NCAA basketball ratings
COOPER is Silver Bulletin’s new Elo‑based power rating for women’s Division I basketball, released March 12, 2026 with ratings through March 11.
dataset Source
saved
The French-YMCA corpus is a new linguistic resource targeting children and adolescents, comprising 39,200 texts and 22,471,898 words drawn from diverse sources with consistent grammar and spelling.
dataset Source
saved
BrowseComp is a 1,266-item OpenAI benchmark (Apr 10, 2025) of intentionally obscure, single-answer browsing tasks created by human trainers with strict checks to ensure difficulty and verifiability.
dataset Source
saved
Humanity’s Last Exam is a 2,500-question test created by scientists to probe deep, cross-disciplinary expert knowledge across mathematics, humanities, natural sciences, ancient languages and niche subfields.
dataset Source
saved
IndQA is a new OpenAI benchmark (released 2025-11-03) of 2,278 culturally grounded, reasoning-focused questions in 12 Indian languages across 10 domains, authored and reviewed by 261 experts.
dataset Source
saved
Qwen 3.5 benchmark comparison website (published 2026-02-25) aggregates verified benchmark scores and head-to-head infographics for models including GPT-5.2, Claude 4.5 Opus, Gemini-3 Pro, Qwen3-Max-Thinking, K2.5-1T-A32B, GPT-OSS-120B and Qwen3.5 variants (397B, 235B, 122B, 35B, 27B).
dataset Source
saved
Epoch AI’s new explorer is a useful primary-source-style attempt to quantify the global installed base of AI accelerators, an increasingly important constraint for model training and deployment.
dataset Source
saved
MTRAG-UN is a 666-task benchmark (released Feb 2026 by Rosenthal et al., IBM) for multi-turn retrieval-augmented generation that emphasizes hard conversational phenomena: unanswerable, underspecified, non‑standalone questions and unclear responses.
dataset Source
saved
SPARTA is a benchmark and prompt-template suite aimed at scalable, principled evaluation of tree-structured multi-hop question answering that must operate over both natural-language text and structured tables.
dataset Source
saved
SumTablets is a curated, machine-readable collection that pairs Unicode cuneiform glyph sequences with scholarly Latin‑alphabet transliterations for 91,606 Sumerian tablets (6.97M glyphs).
dataset Source
saved
PleIAs/SYNTH is a synthetic Q&A/reasoning dataset posted on Hugging Face for text-generation tasks, roughly 3.8k examples and associated with a 0.3B tag and frequent use of the qwen-3-8b-memorization label.
dataset Source
saved
Common Corpus is a two-trillion-token open dataset (announced 2025-06-02) aggregating uncopyrighted or permissively-licensed text and code for LLM pre-training.
dataset Source
saved
Common Corpus is a 2 trillion-token, truly open multilingual dataset released 2026-01-26 (OpenReview id 0wSlFpMsGb) assembled from uncopyrighted or permissively licensed material.
dataset Source
saved
HealthBench is a new OpenAI benchmark (released 2025-05-12) of 5,000 realistic health dialogues with physician-crafted rubrics (48,562 criteria) created by 262 physicians across 60 countries.
dataset Source
saved
PVIR introduces a 95-video benchmark to probe physics-aware video instance removal, focusing on causal side effects (reflections, illumination, shadows) that prior benchmarks ignore.
dataset Source
saved
Free to Use Datasets and Tools to Benchmark Rates and Analyze Affordability was presented as a NARUC CPI innovation webinar (published 2026-03-30) and moderated by Commissioner Christine Needo (appointed February 3, 2024).
dataset Source
saved