Back to registryConditional
Code Review Assistant
Conditionalagt-007
Reviews pull requests for style, common bugs, and missing tests, leaving suggested comments.
- Owner
- Tom Becker
- Developer Experience
- Model
- GPT-5.5
- Data Access
- Internal
- Autonomy
- Suggest
- Value Tier
- Medium
- Hours Saved / mo
- 110
- Connected Tools
- GitHub
Launch Readiness
78/ 100
Rubric breakdown
Seven criteria, each scored 0–4 and weighted. Data Safety and Tool Permission Risk carry hard safety gates.
- Business Value3/4weight ×1.0
- Use Case Clarity4/4weight ×1.0
- Data Safety3/4weight ×2.0
- Tool Permission Risk3/4weight ×2.0
- Evaluation Quality3/4weight ×1.5
- Human Review Coverage3/4weight ×1.5
- Auditability3/4weight ×1.0
Why this verdict
- Aggregate score of 78/100 lands in the Conditional band — solid, but short of the Launch threshold (80).
Evaluation
87%
Accuracy
higher is better
7%
Hallucination
lower is better
9%
Escalation
routed to a human
83%
Satisfaction
higher is better
3.0s
Latency
per task
$0.07
Cost / task
estimated
Failure modes
- Can produce noisy or low-value comments on very large diffs.
- Occasionally misses security-relevant issues in unfamiliar frameworks.
Audit log
| Timestamp (UTC) | Actor | Event |
|---|---|---|
| Jan 22, 2026, 1:00 PM | Tom Becker | Agent registered with comment-only GitHub permissions. |
| Feb 18, 2026, 11:40 AM | Eval Pipeline | Evaluation run completed: 87% useful-comment rate on sampled PRs. |
| Mar 27, 2026, 9:25 AM | Developer Experience | Comment volume tuned down after noisy-feedback signal. |