Mindgard Raises $30M Series A by Turning Attacker Behavioral Intelligence Into Effective AI Defense
Mindgard secured $30 million in Series A funding to expand its automated AI security testing and vulnerability assessment platform.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Mindgard secured $30 million in Series A funding to expand its automated AI security testing and vulnerability assessment platform.
DeepSeek is steeply increasing API prices for its V4 models, closing the cost gap with established major AI rivals.
AI debt accounts for up to 30% of net new credit issuance, exposing bank balance sheets to GPU residual value risks.
OpenAI previews Ultrafast API tier for GPT-5.6 Sol, leveraging Cerebras hardware to deliver up to 750 output tokens per second.
Bank of England research analyzes financial stability risks and market spillovers if US big-tech AI earnings fail to meet valuations.
Standard Chartered argues long-dated capital overcommitment to AI infrastructure poses a greater structural risk to banks than low returns.
Researchers released Backtrader-Bench, an evaluation framework using dynamic MCQs to benchmark LLM coding agents in algorithmic trading.
New benchmark COMPINT reveals LLM context compaction silently drops persistent session constraints and safety instructions in multi-turn sessions.
Research shows LLMs retain latent knowledge after unlearning, which can recover during subsequent fine-tuning or continued training.
Research shows adapting LLMs to specific group preferences increases sycophancy, causing models to prioritize agreement over factual truth.
Research shows semantic rephrasing of benchmark prompts routinely flips LLM output correctness, exposing high evaluation instability.
LabelFusion-TS fuses text models and market time series to classify Federal Reserve communications into hawkish, dovish, or neutral stance.
New research benchmarks the serving costs of agentic memory frameworks against standard context window strategies over long conversations.
New arXiv research benchmarks trustworthiness in small language models, comparing natively trained SLMs against compressed variants.
Researchers released FrontierFinance, a benchmark evaluating AI agents on complex, open-ended investment research workflows.
ArXiv research analyzes performance trade-offs in gist-based context compression, showing where summarization impairs agent memory retrieval.
Research demonstrates China-origin vision-language models systematically reframe sensitive visual inputs rather than issuing explicit refusals.
ToolHazard creates scalable adversarial environments to evaluate LLM agent vulnerability to indirect prompt injections.
Research shows LLM performance rankings flip when generation token budgets vary, with 3–19% of tasks showing accuracy drops at higher caps.
Paper proposes integrating LLM-derived aleatoric and epistemic uncertainty directly into portfolio covariance matrices for small-cap trading.
A systematic study shows LLMs struggle to accurately translate natural language constraints into formal solver code for complex problems.
ProForma-20Q benchmark introduces multi-period financial statement forecasting across 78 line items up to 20 quarters ahead.
Researchers introduce Weightless Fine-Tuning, a decoding-time technique that emulates supervised fine-tuning without updating model weights.
Researchers introduce TradingMoE, a Mixture-of-Experts routing framework designed for LLM-based trading across changing market conditions.
Research demonstrates that post-training quantization for financial time-series models causes accuracy degradation during market regime shifts.
Research introduces test-time scaffolding, using strong models to build runtime harnesses that boost smaller models' task performance.
Paper uses singular learning theory to prove semantic safety constraints lie off-support, meaning data training alone cannot guarantee safety.
ArXiv study reveals agent benchmarks measure task specialization rather than capability, with agent choice causing under 3% of variance.
Research shows that programmatic skill learning for LLM agents optimizes domain adaptation while significantly reducing inference costs.
Paper benchmarks frontier LLMs against native multimodal embedding models like Gemini Embedding 2 on complex text-to-image retrieval.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion