Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs
A systematic study demonstrates that Parameter-Efficient Fine-Tuning (PEFT) on benign datasets degrades LLM safety and fairness alignment.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
A systematic study demonstrates that Parameter-Efficient Fine-Tuning (PEFT) on benign datasets degrades LLM safety and fairness alignment.
Researchers propose a nonparametric distribution regression re-calibration method to prevent overconfident probabilistic predictions.
Researchers developed a tail-aware framework for sub-Weibull data to improve LLM alignment bounds under heavy-tailed reward distributions.
Researchers analyzed alignment faking, where models strategically comply during training to avoid behavioral modification in deployment.
Researchers propose RubricReviewer, a framework using LLMs to generate explicit evaluation rubrics before performing peer reviews.
Researchers introduce AgentMemBench, a standardized benchmark evaluating five distinct long-term memory management strategies for AI agents.
Researchers propose using SLMs trained via SFT and RL as multi-agent routers to improve retrieval quality and reduce orchestration costs.
A new benchmark, XL-DocBench, evaluates LLM performance on evidence-grounded question answering across extra-long, multi-page documents.
A cross-architecture study reveals that domain adaptation of small language models (SLMs) degrades factual calibration and adversarial robustness.
Research reveals that global human evaluation of LLM summaries systematically overlooks individual unfaithful, hallucinated sentences.
Researchers developed PRISMS, a framework using specific MLP neurons to detect and steer LLM tool-use failures before execution.
Research shows ensembling 16 LLMs from 10 families yields only 1.69 distinct semantic perspectives, revealing high redundancy in multi-model architectures.
Academic research demonstrates that LLM progress on difficult tasks reflects overall scaling shifts rather than structural improvements.
Researchers introduce Deep Research Pretraining (DRP), a method to train research agents using offline, naturally occurring link structures.
Researchers propose S4R, a novel method to compress LLM Key-Value caches for long-context windows without high compute or calibration overhead.
Researchers propose a method to prevent cross-lingual collapse in LLMs, enabling native chain-of-thought reasoning in Southeast Asian languages.
Academic paper demonstrates that standard per-chunk filtering in RAG systems fails on multi-hop queries and proposes decomposition.
Researchers introduced OpenART, a framework for safety red-teaming of long-horizon AI agents operating in dynamic, evolving environments.
Researchers propose an on-policy distillation method to address gradient loss in Group Relative Policy Optimization (GRPO) for LLM post-training.
Researchers propose online KV cache compaction techniques for LLM agents to reduce memory bottlenecks in long-context reasoning tasks.
Researchers introduced FinHardBench, a benchmark of 33 tasks evaluating LLMs on generating latency-aware FPGA designs for financial computing.
Researchers proposed ScaleQ-1.58, a post-training quantization framework that preserves reasoning capabilities in 1.58-bit ternary models.
Researchers propose ACE-GraphRAG, an inference-time policy layer designed to dynamically optimize context construction for Hierarchical GraphRAG.
Researchers introduce CrossLex, a benchmark evaluating LLMs on legal reasoning across different jurisdictions using identical facts.
Researchers introduced RH-RAG, a framework designed to improve long-form content generation on locally deployed, open-weight models.
Researchers introduced BiCAA, a framework that applies bidirectional credit assignment to optimize multi-step search agents trained with GRPO.
Researchers evaluate LLMs on their ability to identify economically linked peer firms subject to SEC shadow trading enforcement theories.
Researchers introduce LongChart VQA, a benchmark for evaluating Multimodal LLMs on complex multi-chart reasoning and multi-step inference.
Researchers introduced HopRefusalBench, a benchmark for evaluating how search-augmented LLM agents fail to refuse unanswerable multi-hop queries.
Researchers evaluated LLM performance in goal-directed dialogues across 30 languages, identifying major performance gaps across EU languages.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion