Long-Horizon Forecasting of Complete Financial Statements with Forma
ProForma-20Q benchmark introduces multi-period financial statement forecasting across 78 line items up to 20 quarters ahead.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
ProForma-20Q benchmark introduces multi-period financial statement forecasting across 78 line items up to 20 quarters ahead.
Researchers introduce Weightless Fine-Tuning, a decoding-time technique that emulates supervised fine-tuning without updating model weights.
Researchers introduce TradingMoE, a Mixture-of-Experts routing framework designed for LLM-based trading across changing market conditions.
Research demonstrates that post-training quantization for financial time-series models causes accuracy degradation during market regime shifts.
Research introduces test-time scaffolding, using strong models to build runtime harnesses that boost smaller models' task performance.
Paper uses singular learning theory to prove semantic safety constraints lie off-support, meaning data training alone cannot guarantee safety.
ArXiv study reveals agent benchmarks measure task specialization rather than capability, with agent choice causing under 3% of variance.
Research shows that programmatic skill learning for LLM agents optimizes domain adaptation while significantly reducing inference costs.
Paper benchmarks frontier LLMs against native multimodal embedding models like Gemini Embedding 2 on complex text-to-image retrieval.
Research shows specialist agent decomposition outperforms monolithic LLM prompting in complex European listed real estate financial analysis.
Research identifies bidirectional rationalization in LLM recommendation judges, where zero-shot models convincingly justify opposing outcomes.
Research proposes trajectory-adapted uncertainty quantification for LLM agents to track multi-step error propagation across tool calls.
Researchers applied GRPO reinforcement learning to fine-tune an LLM for financial advice, reportedly beating frontier commercial models.
Paper introduces a hybrid neuro-symbolic framework using LLMs for fact extraction and answer set solvers for deterministic rule reasoning.
Researchers proposed a regime-gated mixture-of-experts model to improve five-day equity realized-volatility forecasting stability.
New research applies speculative decoding to autoregressive time series foundation models, reducing inference latency for long-horizon forecasts.
New arXiv research introduces BrowseSafe, a benchmark and analysis framework evaluating indirect prompt injection risks in AI browser agents.
A new diagnostic stress test measures how system prompts and safety constraints degrade LLM performance relative to standard benchmarks.
Researchers detail a billion-scale Graph Neural Network (GNN) deployed within Weixin Pay for real-time credit fraud detection.
Researchers introduce C-Guard, a framework highlighting that reducing LLM over-refusal of benign prompts silently increases jailbreak vulnerabilities.
Researchers introduce Harness-G, a graph-structured framework to resolve retrieval aliasing in multi-turn search-based RL agents.
Researchers propose a method using almost orthogonal features in language models to perform isolated interventions without downstream entanglement.
Research proposes OrderMoE, a method for deploying Mixture-of-Experts (MoE) models on edge devices by leveraging expert functional similarity.
Research introduces a Stochastic Dimension Zeroth-Order Estimator to improve memory and training efficiency for Physics-Informed Neural Networks (PINNs).
Research reformulates Transformer/Attention mechanisms via measure theory and frequency analysis, claiming hallucination is an inevitable structural LLM limitation.
Research introduces Quantum Port-Hamiltonian Neural Networks (Q-pHNNs) to learn classical dynamics using parameterised quantum circuits.
LakeQuest is a new research benchmark for evaluating question answering systems over heterogeneous, weakly structured data lakes.
Research indicates emotional framing in prompts degrades LLM quantitative reasoning, even when numerical content is identical.
Anthropic is in talks to acquire AI startup Decart for approximately $6 billion in its largest acquisition to date.
Microsoft and S&P Global partner to integrate financial market data and analytics into Microsoft 365 Copilot and agentic workflows.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion