Today for AI

Hacker News AI · 2026/10/9 01:13:08

开源 AI SRE Arena:提供 Kubernetes 故障注入基准,量化评估智能体诊断能力

原标题:Show HN: AI SRE Arena, an Open Benchmark for AI SRE Agents on Kubernetes
68AI 研判分
核心综述

Project Arena 发布了一个厂商中立的开源基准测试框架,专门用于评估 AI SRE 智能体在 Kubernetes 环境中的故障排查能力。该工具支持本地 kind 或 AWS EKS 集群部署,内置 21 种标准故障场景,并通过可配置的评判器对调查记录进行自动化评分。目前已有 Edge Delta、Grafana 和 Claude 等主流方案在该框架下完成了初步对比测试。

报道全文原始报道全文

本文目录5 个章节

Project Arena

A vendor-neutral starter for deploying a disposable Kubernetes fixture, injecting faults, saving investigation records from any product, and scoring completed investigations with a configurable judge. Python 3.10+ is the only Python dependency. Kubernetes operations additionally need kubectl; local cluster creation needs Docker and kind.

Choose local kind or AWS EKS for your cluster, then select the 21-scenario full suite or six-scenario smoke fixture. Cluster type and scenario suite are separate choices.

How it works

流程图

Run commands from the Project Arena directory. Cluster setup, product integration, scenario execution, and scoring are separate steps; follow the sections below in order.

Benchmark results

We compared Edge Delta’s native AI investigations, Grafana’s native AI investigations, and Claude using each platform’s observability CLI across 21 Kubernetes incident scenarios. edx provides access to Edge Delta; gcx provides access to Grafana. All final investigations were evaluated against the same incident facts and scoring rubric using GPT-6-Astra.

Detection and investigation results

Edge Delta detected and investigated 18 scenarios; Grafana detected and investigated 12. Claude was started externally for all 21 scenarios on each platform: 16 alerts and 5 customer reports with edx, and 12 alerts and 9 customer reports with gcx. Detection was not independently measured for Claude, so those cells are shown as —.

Investigation scores use every completed investigation for that column. Implementation readiness excludes cases with no mitigation proposal.

MetricEdge Delta nativeGrafana nativeClaude + edxClaude + gcx
Detection18/21 (85.7%)12/21 (57.1%)——
Root cause analysis15/18 (83.3%)9/12 (75.0%)18/21 (85.7%)19/21 (90.5%)
Blast radius12/18 (66.7%)8/12 (66.7%)18/21 (85.7%)18/21 (85.7%)
Supported final mitigation8/18 (44.4%)5/12 (41.7%)16/21 (76.2%)15/21 (71.4%)
Implementation readiness9/16 (56.2%)5/12 (41.7%)16/21 (76.2%)15/21 (71.4%)

Comparison on the same 12 incidents

This table uses the 12 incident types investigated by both native products, with the corresponding Claude investigations. It controls which scenarios are included, not differences in launch prompts, timing or available evidence.

MetricEdge Delta nativeGrafana nativeClaude + edxClaude + gcx
Root cause analysis11/12 (91.7%)9/12 (75.0%)10/12 (83.3%)10/12 (83.3%)
Blast radius9/12 (75.0%)8/12 (66.7%)10/12 (83.3%)10/12 (83.3%)
Supported final mitigation7/12 (58.3%)5/12 (41.7%)8/12 (66.7%)8/12 (66.7%)
Implementation readiness8/12 (66.7%)5/12 (41.7%)8/12 (66.7%)8/12 (66.7%)

View results for every scenario · Download CSV

What the metrics mean

MetricWhat earns credit
DetectionThe product detected the incident and started an investigation within the observation window.
Root cause analysisThe final report correctly explains what caused the incident.
Blast radiusThe final report correctly identifies the affected workloads and downstream impact, without claiming unsupported outages.
Supported final mitigationThe final recommendation gives a concrete, supported fix or safe containment for the incident, with no remaining incorrect or unsafe advice.
Implementation readinessThe proposal specifies the correction and essential details; normal review, implementation, and rollout checks may remain.

Mitigation and readiness measure proposals, not executed repairs or verified recovery.

See the scoring rubric for grading rules and full results and methodology for verdict breakdowns and evaluation details.

(注:更多技术实现细节与完整 API 文档请查阅原项目 README)