📝text•6 months ago
Evaluation and Regression Gate
Create evaluation runs, submit results, compute metrics, save baselines, and triage regressions before shipping changes.
analysis
⭐1
# Evaluation and Regression Gate
Imported from curated first-party documentation sources.
## What this covers
Use this workflow to formalize benchmarking, regression detection, and baseline management around agent or product changes.
## Use this when
- Validating changes against prior baselines
- Making regressions visible before rollout
- Capturing repeatable evaluation evidence
## Expected outcomes
- Evaluation runs produce comparable metrics
- Baselines are stored after successful validation
- Regression triage becomes a workflow instead of an afterthought
## Source synthesis
- AGENT33/docs/functionality-and-workflows.md (https://github.com/mattmre/AGENT33/blob/main/docs/functionality-and-workflows.md)
- AGENT33/docs/use-cases.md (https://github.com/mattmre/AGENT33/blob/main/docs/use-cases.md)
## Dedupe notes
Combines AGENT33 evaluation lifecycle definitions with evaluation and regression use-case framing.
## Source excerpts
### AGENT33/docs/functionality-and-workflows.md
### 4.3 Evaluation Lifecycle
Flow:
1. Create run (`/v1/evaluations/runs`)
2. Submit task results (`/runs/{id}/results`)
3. Compute metrics + gate report
4. Save baseline (`/runs/{id}/baseline`)
5. Triage/resolve regressions
### AGENT33/docs/use-cases.md
## 4. Evaluation and Regression Gates
Goal:
- Quantify quality and block regressions across PR/merge/release gates.
Use these modules:
- `api/routes/evaluations.py`
- `evaluation/service.py`
- `evaluation/gates.py`
- `evaluation/regression.py`
Typical flow:
1. Create run for gate type (`G-PR`, `G-MRG`, `G-REL`, `G-MON`).
2. Submit task results and quality metadata.
3. Compute metrics and gate verdict.
4. Save baseline for future comparison.
5. Triage/resolve regression records.
Best fit:
- Teams with golden-task style quality gates.
👍0
👁️0
docs