https://marginlab.ai/trackers/claude-code/ marginlab MARGIN LAB HOME BENCHMARK EXPLORERS SWE-Bench Pro Terminal-Bench 2.0 TRACKERS Claude Code Tracker Codex Tracker BLOG CONTACT Last updated: Jan 29, 2026 Claude Code Opus 4.5 Performance Tracker The goal of this tracker is to detect statistically significant degradations in Claude Code with Opus 4.5 performance on SWE tasks. * * Updated daily: Daily benchmarks on a curated subset of SWE-Bench-Pro * * Detect degradation: Statistical testing for degradation detection * * What you see is what you get: We benchmark in Claude Code CLI with the SOTA model (currently Opus 4.5) directly, no custom harnesses. Receive Alerts About Summary Status Degradation Status Shows if any time period has a statistically significant performance drop (p < 0.05). Degradation detected over past 30 days Baseline Baseline Pass Rate Historical average pass rate used as reference for detecting performance changes. 58 % reference rate Daily Pass Rate Daily Pass Rate Percentage of benchmark tasks passed in the most recent day's evaluations. 50 % 50 evaluations 7-day Pass Rate 7-day Pass Rate Aggregate pass rate over the last 7 days. Provides a more stable measure than daily results. 53 % 250 evaluations 30-day Pass Rate 30-day Pass Rate Aggregate pass rate over the last 30 days. Best measure of overall sustained performance. 54 % 655 evaluations Daily Trend Pass rate over time Daily benchmark pass rates over the past 30 days. Hover over legend items for details on each visual element. Pass Rate Daily benchmark pass rate showing the percentage of tasks solved each day. Baseline Historical average pass rate (58%) used as reference for detecting performance changes. Threshold Shaded region around baseline (+-14.0%). Changes within this band are not statistically significant (p >= 0.05). [ ] 95% CI 95% confidence interval for each data point. Toggle checkbox to show/ hide. Wider intervals indicate more uncertainty (fewer samples). Loading... Dashed line at 58% baseline with +-14.0% significance threshold Weekly Trend Aggregated 7-day pass rate Rolling 7-day aggregated pass rates for a smoother trend view with reduced day-to-day noise. Pass Rate 7-day rolling pass rate aggregating daily results for a smoother trend view. Baseline Historical average pass rate (58%) used as reference for detecting performance changes. Threshold Shaded region around baseline (+-5.6%). Changes within this band are not statistically significant (p >= 0.05). [ ] 95% CI 95% confidence interval for each data point. Toggle checkbox to show/ hide. Wider intervals indicate more uncertainty (fewer samples). Loading... Dashed line at 58% baseline with +-5.6% significance threshold Change Overview Performance delta by period Improvement Regression Not Significant 1D Yesterday -8.0% Not Statistically Significant -15% +-14.0% threshold ? With 50 trials, +-14.0% change needed for p < 0.05 +15% 7D Last Week -4.8% Not Statistically Significant -15% +-5.6% threshold ? With 250 trials, +-5.6% change needed for p < 0.05 +15% 30D Last Month -4.1% Statistically Significant -15% +-3.4% threshold ? With 655 trials, +-3.4% change needed for p < 0.05 +15% Methodology The goal of this tracker is to detect statistically significant degradations in Claude Code with Opus 4.5 performance on SWE tasks. We are an independent third party with no affiliation to frontier model providers. Context: In September 2025, Anthropic published a postmortem on Claude degradations. We want to offer a resource to detect such degradations in the future. We run a daily evaluation of Claude Code CLI on a curated, contamination-resistant subset of SWE-Bench-Pro. We always use the latest available Claude Code release and the SOTA model (currently Opus 4.5). Benchmarks run directly in Claude Code without custom harnesses, so results reflect what actual users can expect. This allows us to detect degradation related to both model changes and harness changes. Each daily evaluation runs on N=50 test instances, so daily variability is expected. Weekly and monthly results are aggregated for more reliable estimates. We model tests as Bernoulli random variables and compute 95% confidence intervals around daily, weekly, and monthly pass rates. Statistically significant differences in any of those time horizons are reported. Get notified when degradation is detected We'll email you when we detect a statistically significant performance drop. Email address [ ] Subscribe Thanks for subscribing! Check your email to confirm. Something went wrong. Please try again. marginlab MARGIN LAB Home Blog (c) 2026 Marginlab. All rights reserved.