XALPHA: A Memory-Driven AI Quant Researcher for Hypothesis-to-Code Alpha Discovery
XALPHA proposes a memory-driven AI quant researcher for end-to-end alpha discovery from hypothesis to code, leveraging LLMs.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
XALPHA proposes a memory-driven AI quant researcher for end-to-end alpha discovery from hypothesis to code, leveraging LLMs.
Research views self-attention in Transformers as a connection walk, identifying it as a specific operator on token-position graphs.
Soofi S 30B-A3B, a sovereign, open-source MoE hybrid Mamba Transformer model for German and English, claims high throughput for long context.
SpurLens detects spurious correlations in Multimodal LLMs using GPT-4 and object detectors, revealing inherent biases.
Research paper argues LLM-based social simulations need boundaries due to their tendency to produce homogeneous, 'average persona' outputs, limiting diversity.
Research introduces a constraint-aware hierarchical search method for fine-grained classification tasks driven by regulatory rules, not just semantic similarity.
Research compares Vision-Language Models (VLMs) across domains, highlighting that benchmark performance does not predict real-world behavior across varied datasets.
New benchmark SWE-MERA aims to dynamically evaluate LLMs on software engineering tasks, addressing data contamination issues in SWE-bench.
CRINN introduces a contrastive reinforcement learning method to optimize Approximate Nearest-Neighbor Search (ANNS) algorithms, targeting execution speed.
Nested-ReFT proposes an efficient reinforcement learning method for large language model fine-tuning (ReFT) using off-policy rollouts to improve reasoning.
Research addresses emotion recognition in signers, creating new datasets for Japanese Sign Language and British Sign Language to overcome data scarcity and grammatical-affective expression overlap.
TagSpeech presents a unified LLM-based framework for end-to-end multi-speaker ASR and diarization with fine-grained temporal grounding.
MUGEN benchmark reveals large audio-language models (LALMs) struggle with multi-audio understanding, especially with increased concurrent inputs.
Research introduces RecursiveMAS, a framework extending recursive computation from single LLMs to multi-agent systems for deepened reasoning.
Gefen is a new memory-efficient optimizer for deep learning, reducing AdamW's memory footprint by sharing and quantizing moment estimates.
Research demonstrates a multi-agent LLM framework for automated generation of verifiable reaction rules, reducing manual encoding for complex systems.
Research explores how agent memory characteristics influence conceptual alignment and shared meaning emergence in non-partnership coordination games.
Research introduces ARMOR, a method to stabilize LLM reinforcement learning (RL) by using off-policy anchor samples to prevent over-optimization.
ANCHOR is a new automated auditing framework stress-testing CLI agents on illegal tasks derived from public US court cases.
GigaAM Multilingual is a foundation model for automatic speech recognition focused on underrepresented Central Asian languages, pre-trained on 2M hours of audio.
Researchers introduced ChartSync, a benchmark to evaluate visuo-logical cascading editing in generative models for structured statistical charts.
Research explores detecting confident hallucinations in LLMs for financial Q&A by analyzing internal model activations, beyond observable outputs.
Frontier LLMs achieved only 68% accuracy in an agentic clinical reasoning framework requiring proactive data requests, exposing information-seeking failures.
Research proposes a method to fingerprint and verify LLMs from single-token outputs, addressing model deviations in opaque serving chains.
New research proposes a grammar-driven approach to code watermarking, aiming to improve detectability without sacrificing code quality.
Research proposes an AI framework to identify expired patents, analyze technology trends, and translate disclosures into business commercialization pathways.
A new survey details the theoretical foundations and deployment challenges of LLM watermarking for provenance and misuse detection.
Researchers introduced MAG, a benchmark and harness for multimodal web-agents to generate actions and guides across multi-page web tasks.
Research indicates distinct geometric fingerprints in transformer hidden states from identity-specifying prompts across different model training regimes.
Research identifies a "self-correction blind spot" in autoregressive LLMs, where they fail to fix their own errors despite correcting external ones.
Research shows length penalties for chain-of-thought can obscure a model's underlying reasoning influences, making monitoring harder despite accuracy.
EvoCUA-1.5 presents online reinforcement learning for multi-turn computer-use agents, addressing limitations of static training data.
Research explores optimal placement and cost of 'oracle' correctors to guide unreliable multi-agent swarms toward correct consensus.
Research proposes a dynamical systems model to interpret latent Chain-of-Thought (CoT) reasoning in LLMs, addressing current interpretability gaps.
Research introduces Low-Rank Attention Residuals (LR-AttnRes) to optimize LLM architecture by decoupling routing from representation, using lower-dimensional keys.
Research demonstrates a method to detect if an LLM was distilled from another model, particularly in a reference-based setting.
Research evaluates LLM-generated English grammar drills for EFL students, analyzing question types, cognitive load, and CEFR alignment against student performance.
Research paper explores metacognition in LLMs, its current limitations, and future opportunities for more capable and transparent AI systems.
New benchmark, AdvancedMathBench, evaluates LLMs on advanced mathematical proof generation and verification beyond olympiad-style problems.
Research explores how 'temperature' parameter in RAG systems can transmit, amplify, or suppress ideological biases from retrieved content.
New research proposes Theory-Grounded and Culture-Aware Multilingual Moral Reasoning (MET) for language models, addressing English-centric biases.
JobHop v2, a large-scale career trajectory dataset from unstructured resumes, is released to support workforce planning and labor market analysis.
Research explores if LLMs exhibit an asymmetry between language production and perception, similar to human psycholinguistics, using token probability.
RAGU is a new open-source GraphRAG engine that improves knowledge graph construction and retrieval through two-stage extraction and deduplication.
Research introduces Associative Recurrent Memory Transformer (ARMT) to extend LLM context length beyond quadratic scaling, improving efficiency.
Research exposes limitations in benchmark-trained models for Bangla hate speech detection, citing cultural context and linguistic nuances.
Research examines LLMs' political biases, finding models are predominantly left-leaning, mimicking training data, but can be steered.
Research proposes UMoE, a method to re-optimize Mixture-of-Experts (MoE) models during domain-specific fine-tuning, improving expert utility.
Research proposes a multimodal RLHF framework for direct image-to-modern Vietnamese translation of degraded Han-Nom manuscripts.
Researchers introduced ToFu, a white-box, token-efficient agent harness for LLMs that reads codebases, edits files, and runs commands.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion