Pangram 4 Technical Report
Pangram Labs released a technical report for Pangram 4, claiming a 0.9916 AUROC and a 0.0041% false positive rate for AI text detection.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Pangram Labs released a technical report for Pangram 4, claiming a 0.9916 AUROC and a 0.0041% false positive rate for AI text detection.
Mercor and Ramp introduced APEX-Accounting, a private benchmark of 160 multi-file tasks designed to evaluate LLMs on accounting workflows.
Researchers analyzed GPT-4-Turbo to identify and measure implicit bias toward people with intellectual disabilities in LLM-generated stories.
Researchers introduce GPT-Red, an automated self-play red-teaming agent designed to find prompt injection vulnerabilities in frontier models.
Researchers find RL-trained reasoning models develop superior internal representational quality over SFT models for math problem-solving.
Research identifies 'trust inflation' in LLM evaluation, where aggregating weak and strong metrics masks vulnerabilities in model performance.
Researchers identify a 'confounder trap' in causal inference from text, where representations learned to control for confounding encode treatment.
Researchers proposed IRIS, a framework that extracts dynamic user personas from implicit interaction streams rather than explicit feedback.
Researchers demonstrate that audio-capable LLMs can be jailbroken solely through variations in vocal delivery (prosody) with fixed transcripts.
Researchers introduce SecRespond, a benchmark specifically designed to evaluate LLM agents in real-world post-compromise incident response tasks.
Research shows frontier multimodal models generate structured, biased confabulations based on demographic descriptors when input images are missing.
Researchers introduced Setoka, a new benchmark designed to evaluate personalized AI agents on retrieving and inferring user characteristics.
Researchers propose an on-policy distillation routing method to protect fine-tuned LLMs from malicious safety-realignment bypasses.
Researchers define 'linguistic monoculture,' showing that widespread LLM-assisted drafting reduces population-level variations in text.
Researchers introduce MindForge, a framework training small language models in whole-life-cycle software generation via synthetic data.
Researchers introduce SpecFirst, an agentic framework designed to improve program synthesis from scratch using behavioral exploration.
A study of Gemini 2.0 Flash and ChatGPT-4o finds significant diagnostic inconsistency under rephrased prompts and irrelevant content.
Researchers introduce ML2B, a benchmark of 35 Kaggle competitions in 14 languages to evaluate LLMs on cross-lingual ML pipeline generation.
Researchers propose ARC-Encoder, a context compression technique that works without modifying or fine-tuning the target LLM architecture.
Researchers analyzed how context alters truth representation vectors inside LLM activations, finding context deforms internal truth geometry.
Researchers introduced DialectLLM, a framework to generate multi-dialectal conversational data to improve LLM performance on non-American English.
Researchers introduced CustomerSim, a benchmark evaluating multimodal language models on their ability to simulate persona-driven customer behavior.
Researchers propose SkillTTA, a method that synthesizes temporary skills at test-time to improve LLM agent task performance.
StateRAG introduces typed state contracts to manage complex retrieval-augmented generation pathways, externalizing retrieval decisions from model contexts.
Researchers propose SkillCAT, a framework to improve LLM agent skill self-evolution by filtering low-quality trajectory edits and context bloat.
Researchers identify 'F1 Inflation' in LLM evaluation, showing that prompt framing and numeric anchoring distort error-detection metrics.
An academic study on arXiv evaluates whether scaling compute in LLMs inherently improves the fidelity of social and behavioral simulations.
Researchers developed SAR, a LoRA-based adapter designed to force fine-tuned language models to self-report hidden or backdoored behaviors.
Researchers identify 'library drift' in self-evolving LLM skill libraries, showing unmanaged skill accumulation degrades retrieval and stalls performance.
Microsoft committed $130 billion to data center leases, driving its intelligent cloud segment revenue up 32% to $39.3 billion.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion