Do Generative Models Keep Time? A Time-Aware Evaluation of Synthetic Sequential Tabular Data
Research details a new evaluation framework for synthetic sequential tabular data, focusing on time-awareness to prevent illogical time sequences.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research details a new evaluation framework for synthetic sequential tabular data, focusing on time-awareness to prevent illogical time sequences.
Research systematically investigates memorization behaviors in Rectified Flow generative models, focusing on generalization vs. data recall.
Research introduces an 'Attention-Discounted Adaptive Sampler' for masked diffusion language models to improve parallel token generation in inference.
Researchers propose PRISM Edit, a new LLM editing method allowing temporal facts to be updated while retaining historical accuracy without full retraining.
New research proposes a constrained two-view framework for Graph Neural Networks (GNNs) to improve node prediction accuracy by decoupling feature transformation and neighborhood aggregation, addressing topology noise and heterophily.
Research addresses runtime safety for learned sUAS separation policies under degraded GNSS, highlighting fundamental architectural questions for deployment.
A new benchmark, Loci Similes, is proposed for evaluating language models' ability to extract intertextual links in Latin literature.
Research paper details an LLM-powered pipeline for automated data extraction and structuring from scientific literature, exemplified with concrete materials.
Research paper models benchmark hacking in ML contests, showing how models are tuned to score highly without true generalization.
Research identifies novel 'function hijacking' attacks against agentic LLMs, exploiting vulnerabilities in external function calling mechanisms.
Research surveys dynamic model routing and cascading strategies for LLM inference to optimize performance and cost by selecting models based on query complexity.
REALM proposes fine-tuning LLMs with noisy human annotations by jointly learning model parameters and annotator reliability, surpassing standard aggregation.
Research suggests LLM-generated labels can rival human labels in active learning for hostility detection, potentially reducing annotation costs.
Research proposes a novel method, GRADE, using gradient subspace dynamics to probe LLM internal knowledge gaps, aiming for better confidence detection.
Research evaluates large language models' effectiveness in generating multilingual synthetic data for training smaller models, highlighting capability gaps in non-English languages.
Research on supervised uncertainty quantification for LLMs finds existing probe methods are not robust under distribution shift, impacting hallucination detection.
Researchers augmented a deep anomaly detection dataset for batch distillation with simulation data to improve model training for industrial processes.
Scalable Capital opened its investment platform to external AI assistants, allowing direct integration with user brokerage accounts.
AWS Agent Registry is now generally available, providing a governed catalog for managing and discovering agents, tools, and skills across organizations.
OpenAI is reportedly offering select enterprise clients outcome-based pricing, charging only when models perform specified tasks successfully.
FSB Chair Andrew Bailey warns global regulators that autonomous frontier AI models pose systemic cyber and operational financial risks.
Bank of England Governor Andrew Bailey warned that frontier AI models pose severe cyber security risks to global financial stability.
FSB Chair Andrew Bailey warned G20 leaders that unmonitored frontier AI releases pose systemic risks to global financial stability.
FSB Chair Andrew Bailey warned G20 finance ministers and central bankers about market vulnerabilities and systemic risks from frontier AI.
Research finds computer-use agents fail silently 90% of the time despite high overall benchmark task scores, proposing a runtime oversight gate.
Researchers proposed a zk-SNARK framework using adversarial probes to detect post-deployment drift in proprietary third-party LLMs.
Defense training against prompt injections systematically degrades LLM agents' ability to perform multi-step autonomous tool executions.
Research shows constrained generation can silently suppress tool calls, causing models to hallucinate extractions while passing fidelity checks.
Research proves mathematical differential privacy bounds do not prevent adaptive extraction of memorized training data in LLMs.
Research demonstrates LLM safety refusal behavior varies across random seeds and temperatures, invalidating single-shot safety evaluations.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion