Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA
Research on Portugal's open-weight national language model, AMALIA, evaluates its trustworthiness as a 'scientific instrument' for community measurement.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research on Portugal's open-weight national language model, AMALIA, evaluates its trustworthiness as a 'scientific instrument' for community measurement.
Research shows a 146M-parameter small language model can detect LLM behavioral issues like sycophancy and confabulation more effectively than human raters.
Research introduces a three-stage calibration framework to evaluate how post-training methods (SFT, RL, OPD) reshape LLM confidence during Chain-of-Thought reasoning.
Researchers introduced Unsupervised Multimodal Clustering (UMC) for semantics discovery in multimodal utterances without pre-existing labels.
Research shows Vision Language Models (VLMs) can 'see' non-existent visual illusions, highlighting limitations in their perceptual reasoning.
Research introduces "Mediator" for memory-efficient LLM merging, reducing parameter conflicts and improving performance over traditional averaging.
Research introduces Prismatic Synthesis, a gradient-based method for generating diverse training data to improve LLM generalization in reasoning tasks.
MMR-V, a new benchmark, addresses MLLM limitations in multi-frame evidence location and multimodal reasoning for video data.
Researchers introduced LEGO Co-builder, a hybrid benchmark for fine-grained vision-language models focused on multimodal assembly instructions.
Research uses LLMs to identify social biases against homelessness in online text and city council discourse, highlighting systemic issues.
Research describes DisarmRAG, a retriever-centric poisoning attack disabling LLM self-correction in Retrieval-Augmented Generation systems.
Researchers introduced ChipChat, a low-latency cascaded conversational agent architecture in MLX, aiming for real-time on-device voice agents.
Research introduces Sequential Tool Attack Chaining (STAC), a multi-turn attack framework exploiting LLM agent tool use to chain harmless tools for harmful ops.
Research proposes Parallel Decoder Transformer, a model-intrinsic architecture for generating multiple document sections concurrently from a single LLM.
Research introduces GradAlign, a method to improve LLM reinforcement learning performance by selecting high-quality training problems, addressing RL's non-stationarity.
FormulaCode introduces a new benchmark to evaluate LLM coding agents' ability to optimize entire codebases, moving beyond synthetic, single-objective tasks.
A new LLM-based indicator developed by PLOS and DataSeer measures research data reuse in scholarly publications, showing a 43% reuse rate.
Research paper argues current coding benchmarks are misaligned with agentic software engineering due to collapsed scoring and lack of component-level signal.
OmniAgent is a research proposal for a new omni-modal agent architecture designed for efficient long video understanding through active perception.
Theoria introduces a verification architecture for AI system answers, converting solutions into auditable, typed state transitions to bridge formal proof certainty with LLM coverage.
FOI-O proposes a global ontology and verification framework for modeling Freedom of Information (FOI) processes using public records.
Research explores moving memory inside the LLM agent's observation-reasoning-act loop, allowing constant read/write to address latency challenges.
RuBench 1.0 is a new benchmark for evaluating agentic coding models using 25 real-world repository-level tasks with native Russian specifications.
NEXTDC, an Australian data center operator, secured new customer contracts, increasing its contracted utilization by 11% in Q2.
China's GigaAI (Jijia Vision) is reportedly planning a Hong Kong IPO by 2026, joining other Chinese AI firms seeking public debuts.
Anthropic's $1.5B copyright settlement is approved, resolving a specific case but leaving broader training data legal issues open.
David Vélez (Nubank CEO) and Robin Vince (BofA, former Goldman Sachs CFO) join the OpenAI Foundation and OpenAI Group PBC boards.
Apple ML Research proposes an environment-free synthetic data generation method for training API-calling LLM agents, using LLMs as digital world models.
Apple ML Research proposes Calibrated Sparse Attention to speed up text-to-video generation in diffusion models by skipping negligible computations.
South Korea's early July exports increased, driven by strong global demand for semiconductors, especially those tied to AI.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion