Today for AI

Hacker News AI · 10/9/2026, 01:13:08

Open-Source AI SRE Arena: Benchmarking Kubernetes Fault Diagnosis Agents with Standardized Scoring

68AI Score
Executive Summary

Project Arena released a vendor-neutral open-source benchmark framework designed to evaluate AI SRE agents' fault diagnosis capabilities in Kubernetes environments. Supporting local kind or AWS EKS clusters, it includes 21 standardized fault scenarios and uses a configurable judge to automatically score investigation records. Preliminary comparisons have already been conducted for solutions like Edge Delta, Grafana, and Claude using this framework.

SOURCE COVERAGEOriginal coverage

Contents5 sections

Project Arena

A vendor-neutral starter for deploying a disposable Kubernetes fixture, injecting faults, saving investigation records from any product, and scoring completed investigations with a configurable judge. Python 3.10+ is the only Python dependency. Kubernetes operations additionally need kubectl; local cluster creation needs Docker and kind.

Choose local kind or AWS EKS for your cluster, then select the 21-scenario full suite or six-scenario smoke fixture. Cluster type and scenario suite are separate choices.

How it works

Diagram

Run commands from the Project Arena directory. Cluster setup, product integration, scenario execution, and scoring are separate steps; follow the sections below in order.

Benchmark results

We compared Edge Delta’s native AI investigations, Grafana’s native AI investigations, and Claude using each platform’s observability CLI across 21 Kubernetes incident scenarios. edx provides access to Edge Delta; gcx provides access to Grafana. All final investigations were evaluated against the same incident facts and scoring rubric using GPT-6-Astra.

Detection and investigation results

Edge Delta detected and investigated 18 scenarios; Grafana detected and investigated 12. Claude was started externally for all 21 scenarios on each platform: 16 alerts and 5 customer reports with edx, and 12 alerts and 9 customer reports with gcx. Detection was not independently measured for Claude, so those cells are shown as —.

Investigation scores use every completed investigation for that column. Implementation readiness excludes cases with no mitigation proposal.

MetricEdge Delta nativeGrafana nativeClaude + edxClaude + gcx
Detection18/21 (85.7%)12/21 (57.1%)——
Root cause analysis15/18 (83.3%)9/12 (75.0%)18/21 (85.7%)19/21 (90.5%)
Blast radius12/18 (66.7%)8/12 (66.7%)18/21 (85.7%)18/21 (85.7%)
Supported final mitigation8/18 (44.4%)5/12 (41.7%)16/21 (76.2%)15/21 (71.4%)
Implementation readiness9/16 (56.2%)5/12 (41.7%)16/21 (76.2%)15/21 (71.4%)

Comparison on the same 12 incidents

This table uses the 12 incident types investigated by both native products, with the corresponding Claude investigations. It controls which scenarios are included, not differences in launch prompts, timing or available evidence.

MetricEdge Delta nativeGrafana nativeClaude + edxClaude + gcx
Root cause analysis11/12 (91.7%)9/12 (75.0%)10/12 (83.3%)10/12 (83.3%)
Blast radius9/12 (75.0%)8/12 (66.7%)10/12 (83.3%)10/12 (83.3%)
Supported final mitigation7/12 (58.3%)5/12 (41.7%)8/12 (66.7%)8/12 (66.7%)
Implementation readiness8/12 (66.7%)5/12 (41.7%)8/12 (66.7%)8/12 (66.7%)

View results for every scenario · Download CSV

What the metrics mean

MetricWhat earns credit
DetectionThe product detected the incident and started an investigation within the observation window.
Root cause analysisThe final report correctly explains what caused the incident.
Blast radiusThe final report correctly identifies the affected workloads and downstream impact, without claiming unsupported outages.
Supported final mitigationThe final recommendation gives a concrete, supported fix or safe containment for the incident, with no remaining incorrect or unsafe advice.
Implementation readinessThe proposal specifies the correction and essential details; normal review, implementation, and rollout checks may remain.

Mitigation and readiness measure proposals, not executed repairs or verified recovery.

See the scoring rubric for grading rules and full results and methodology for verdict breakdowns and evaluation details.

(注:更多技术实现细节与完整 API 文档请查阅原项目 README)