MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams
MIRA-Ev is a new clinical NLP benchmark for granular evidence detection and relational reasoning, using Spanish medical licensing exams.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
MIRA-Ev is a new clinical NLP benchmark for granular evidence detection and relational reasoning, using Spanish medical licensing exams.
Research proposes RLAES, an LLM framework using reinforcement learning with rubric rewards for automated essay scoring and feedback generation.
Research explores cost-quality tradeoffs in Reinforcement Learning with Verifiable Rewards (RLVR) for Neural Machine Translation, particularly for legal documents.
Replicated prior research on human label variation in Natural Language Inference (NLI), confirming lower agreement for non-upward monotonicity operators.
Research evaluates multimodal LLMs' theory-of-mind (ToM) reasoning in multi-party meetings, identifying current limitations beyond overt signals.
Research investigates mitigating cross-lingual factual inconsistency in LLMs at inference time, addressing biases toward high-resource languages.
Research experimentally evaluates prompt design factors—format, instruction count, and context length—on LLM instruction adherence and hallucination.
New benchmark, GAMUT, focuses on evaluating the factual completeness of open-ended LLM generations, addressing the gap beyond factual precision.
Research identifies "repetitive copying" as a critical failure mode in long-context LLM reasoning, proposing evidence-aware reinforcement learning.
MUX proposes distilling discrete LLM reasoning steps into continuous multiplexed tokens for more efficient, high-bandwidth computation.
Research evaluates state compression in two-agent LLM relays, focusing on how information bottlenecks affect constraint preservation in planning tasks.
Research identifies a 'hidden auto-regressive risk regime' where large language models' answers degrade faster, compounding mistakes.
Research proposes Multi-Task On-Policy Distillation (MTOPD) using soft prompts for LLM self-distillation, avoiding weight drift and input post-hoc rationalization.
Research on multi-step, tool-augmented LLM agents identifies 'binding drift' where agents silently misattribute entities across sequential actions.
EmoEUS proposes an explicit uncertainty supervision framework for multimodal emotion recognition in conversation, addressing modality-specific uncertainty.
Research indicates transformers without feed-forward networks (attention-only) can match or exceed standard transformers in performance when parameters are controlled.
Research identifies 'Safety Drift' and 'Operational Hallucination' in multi-turn LLM agent interactions, degrading initial alignment over time.
Research explores using LLMs for multilingual privacy policy analysis, extending beyond English-centric systems to improve transparency audits.
AutoIndex is a framework for learning representation programs that transform raw documents for retrieval systems, optimizing document preprocessing.
Researchers introduced PLAID-PRF, a method for pseudo-relevance feedback over PLAID, improving dense retrieval effectiveness for fine-grained token-level interactions.
Research proposes "token inoculation" to safely retain dual-use knowledge in LLMs, gating hazardous content with control tokens instead of erasing it.
Researchers demonstrated a compact Hindi text-to-speech model by distilling a large flow-matching teacher model (IndicF5, 337M parameters) with limited data.
Research explores using 'semantic primes' as non-circular explanations for emotion mechanisms in large language models, addressing current limitations.
Research finds non-invasive EEG-to-text (EEG2Text) models fail to generate meaningful decoding without teacher-forcing evaluation.
AgentDebugX is an open-source toolkit designed to improve failure observability, attribution, and recovery for LLM agents by organizing debugging as a closed loop.
AFIR, a Romanian government agency, deployed RAGAL, a fully local RAG-based assistant for technical support, adhering to zero data egress.
Research explores using bounding boxes to crop student responses, improving Small Language Model (SLM) performance on vision-based grading tasks.
Research proposes using AI trademark data to complement patent data, providing a new lens on AI innovation development and diffusion beyond technical invention.
Research proposes AI Tour Meeting, a framework using multiple LLM agents with distinct personas to collaboratively plan group travel through natural language discussion.
Researchers propose Hybrid Hindsight Self-Distillation (H²SD), an RLVR method addressing sparse supervision in LLM reasoning tasks.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion