| 31 Jul 2026 | arXiv cs.LG — Machine Learning | Can Deep Generative Models Reproduce Non-Stationary Gaussian Random Fields? ↗ | Investigate | Next 12 months | deep generative models, spatial modeling, model evaluation |
| 31 Jul 2026 | arXiv cs.LG — Machine Learning | Beyond the Bidirectional Promise: Re-evaluating the Robustness of Diffusion Language Models ↗ | Investigate | Next 12 months | model evaluation, llm security, model robustness |
| 31 Jul 2026 | arXiv cs.LG — Machine Learning | What Is The Performance Ceiling of My Classifier? Utilizing Category-Wise Influence Functions for Pareto Frontier Analysis ↗ | Investigate | Next 12 months | model evaluation, data quality, model performance |
| 31 Jul 2026 | arXiv cs.LG — Machine Learning | Dynamically Scaled Activation Steering ↗ | Investigate | Next 12 months | safety alignment, responsible ai, model evaluation |
| 31 Jul 2026 | arXiv cs.LG — Machine Learning | Uncertainty quantification for trustworthy deep learning: Methods and measures ↗ | Investigate | Next 12 months | uncertainty quantification, model risk, explainability |
| 31 Jul 2026 | arXiv cs.LG — Machine Learning | Transporting Task Vectors across Different Architectures without Training ↗ | Investigate | Next 12 months | model adaptation, fine tuning, model optimization |
| 29 Jul 2026 | arXiv cs.CL — Computation and Language | Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models ↗ | Investigate | Next 12 months | llm security, model evaluation, safety alignment |
| 29 Jul 2026 | arXiv cs.CL — Computation and Language | Evaluation of Adversarial Robustness in Arabic Language Models ↗ | Investigate | Next 12 months | llm security, model evaluation, responsible ai |
| 29 Jul 2026 | arXiv cs.CL — Computation and Language | Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models ↗ | Investigate | Next 12 months | llm security, model governance, intellectual property |
| 29 Jul 2026 | arXiv cs.CL — Computation and Language | VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation ↗ | Investigate | Next 12 months | multimodal reasoning, rag, visual llm |
| 29 Jul 2026 | arXiv cs.CL — Computation and Language | Contrastive Weak-to-strong Generalization ↗ | Monitor | Next 12 months | model training, llm scaling, safety alignment |
| 29 Jul 2026 | arXiv cs.CL — Computation and Language | WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing ↗ | Investigate | Next 12 months | agentic ai, model evaluation, enterprise deployment |