DeepSeek's V4.1-Flash Beats Claude Opus 5 On Agent Benchmarks At A Fraction Of The Price

DeepSeek's V4.1-Flash Beats Claude Opus 5 On Agent Benchmarks At A Fraction Of The Price

0:00 / 0:56
News

DeepSeek's V4.1-Flash Beats Claude Opus 5 On Agent Benchmarks At A Fraction Of The Price

calendar_today Date:
schedule Duration: 0:56
database
Summary Report

DeepSeek-V4.1-Flash edges Claude Opus 5 on DeepSWE and CyberGym, ships native vision and a 1M context under an MIT licence, and costs a fraction of the frontier labs' prices.

  • 01. 74.2 on DeepSWE vs Opus 5's 74.0; 88.1 on CyberGym vs 84.5 for Opus 5 and GPT-5.6 Sol; Opus still wins Terminal-Bench 43.3 to 30.0
  • 02. Off-peak API pricing of $0.15 per million uncached input tokens and $0.60 per million output; DeepSeek says it beats its own larger V4-Pro on several benchmarks
DeepSeek has released V4.1-Flash, the smallest model in its new causal encoder-decoder architecture family: 552B parameters with 8B active in prefill and 16B in decode, native image understanding, a one-million-token context and MIT-licensed weights. It scores 74.2 on DeepSWE against Claude Opus 5's 74.0 and 88.1 on CyberGym against 84.5, at $0.15 per million input tokens and $0.60 per million output off-peak.