Eval

Evaluate and rank agent results by metric or LLM judge for an AgentHub session.

Development / Engineeringdevelopmentengineering
by AgentVoltv1.0.0Published 1y ago
Free to sign up · every skill included with AgentVolt Pro

About this skill


name: eval description: Use when Evaluate and rank agent results by metric or LLM judge for an AgentHub session.

Eval

Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents. — One of 337+ skills in Ali Reza's multi-agent claude-skills collection (Claude, Codex, Gemini, Hermes; ~19k stars).

What you get

  • Public GitHub repo (alirezarezvani/claude-skills)
  • the eval skill folder with SKILL.md. Part of a 337-skill / 30-agent / 70-command install.

Customize your output

  • Fork the repo and adapt the skill's instructions and references to your workflow.

Example output

Activates automatically when your request matches Eval; chains with the other skills, agents, and commands in the collection.

Best for

Creators, builders, and teams using Claude Code.

SKILL.md preview

SKILL.md
---
name: eval
description: Use when Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
---

# Eval

Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents. — One of 337+ skills in Ali Reza's multi-agent claude-skills collection (Claude, Codex, Gemini, Hermes; ~19k stars).

## What you get

- Public GitHub repo (alirezarezvani/claude-skills)
- the eval skill folder with SKILL.md. Part of a 337-skill / 30-agent / 70-command install.

## Customize your output

… (sign up to view the full skill)
Sign up to view, copy, and install the full skill