Skip to content

Repository files navigation

Agentic Eval Flow

Automated Tekton-orchestrated pipeline on OpenShift for evaluating AI artifacts:

  • Skills -- Measures skill efficacy by comparing agent performance with and without skills (A/B "gap" testing)
  • MCP Servers -- Validates MCP server implementations via task-based verification
  • Agents -- Evaluates full agent behavior using Harbor (general agents) or A2A protocol (A2A-compliant agents)

Produces statistical reports with pass rates, uplift metrics, significance tests, and a unified scorecard.

Pipelines

Pipeline Purpose Key Differences
CI Pipeline Full evaluation for new submissions Includes security scan, quality review, artifact generation
Monitoring Pipeline Regression detection for deployed artifacts Includes degradation check against historical baseline, Slack alerts

How It Works

The pipeline executes in six main stages, with engine-specific steps within each:

1. Prepare

  • Clone submission repository
  • Validate structure and metadata.yaml schema
  • AI-assisted generation of missing test artifacts (optional)

2. Test (CI Pipeline only)

  • Quality Review -- AI-powered review of skill/test coherence (advisory)
  • Security Scan -- Cisco AI Defense scan for prompt injection, data exfiltration risks
  • Security & Quality Scan -- harness-eval deterministic scan (27 rule categories covering prompt injection, credential access, obfuscation, coercive overrides, stealth persistence, data exfiltration, description quality, broken references, and more)

3. Evaluate

Six evaluation engines, each suited for different artifact types:

Engine Evaluates Comparison Mode Container Isolation
Harbor Skills, general agents A/B (treatment vs control) Yes
ASE Skills only A/B (treatment vs control) No
A2A A2A-protocol agents A/B (treatment vs control) Yes
MCPChecker MCP servers Single-agent task verification No
AEH Agents, skills Judge-based evaluation Yes (K8s Harbor pods)
AEH OpenShell OpenClaw OpenClaw agents via forge-saw Judge-based (AEH scoring) Yes (SAW VM + inner sandbox)

Engines are implemented in abevalflow/engines/ using a registry pattern.

4. Analyze

  • Compute pass rates, uplift (gap), statistical significance (p-value)
  • Generate report.json and report.md
  • Aggregate gate results into unified scorecard.json
  • Monitoring only: Check for degradation against historical baseline

5. Store

  • Upload reports and artifacts to MinIO
  • Record results to PostgreSQL for historical analysis

6. Cleanup

  • Remove temporary workspaces and artifacts

Configuration

Submission routing and gate configuration are defined in metadata.yaml. For AEH, eval.yaml defines execution, the dataset and judges; PipelineRun parameters and installed Tekton resources define infrastructure and source revisions:

name: my-submission
eval_engine: harbor              # harbor, ase, a2a, mcpchecker, aeh, or aeh_openshell_openclaw
persona: general                 # Agent persona for Harbor/A2A

experiment:
  n_trials: 20                   # Number of evaluation attempts

gate_policy:
  default_mode: warn
  combination: all_pass
  gates:
    evaluation:
      mode: block
      threshold: 0.0
    security:
      mode: warn

See Gate Policy Configuration for full options.

Repository Structure

agentic_eval_flow/
├── Docs/                    # ADR, implementation plan, guides
├── pipeline/
│   ├── pipeline.yaml        # Main pipeline definition
│   ├── triggers/            # EventListener, TriggerTemplate, TriggerBinding
│   └── tasks/               # Tekton task definitions (phases, components, post)
├── templates/               # Jinja2 templates (Dockerfiles, test.sh, task.toml)
├── scripts/                 # Python scripts invoked by pipeline tasks
├── config/                  # K8s manifests (RBAC, PostgreSQL, LiteLLM)
└── tests/                   # Unit and integration tests

Related Repositories

Repository Purpose
skill-submissions Submission intake -- users push skills, MCP evals, and agent evals here
laude-institute/harbor Harbor upstream (PyPI harbor==0.20.0) -- classic A/B uses stock Harbor + OpenShift custom env
opendatahub-io/agent-eval-harness AEH -- KubernetesEnvironment base for the OpenShift custom env plugin
cisco-ai-defense/skill-scanner Security scanner for prompt injection and data exfiltration detection
harness-eval Deterministic security and quality scanner for skill submissions (27 rule categories, 97 rules)

LLM Access

The pipeline is LLM-agnostic. Three modes are supported:

Mode Proxy Required?
Direct API key (Anthropic, OpenAI, etc.) No
opencode + self-hosted model (vLLM, Ollama) No
Google Vertex AI + LiteLLM proxy Yes

Prerequisites

  • OpenShift cluster with Pipelines operator (Tekton)
  • For Harbor/build-based flows: a container registry with push credentials and the Harbor fork with OpenShift backend
  • For AEH OpenShell: an existing Forge SAW deployment and the namespace setup below
  • LLM access (one of the three modes above)
  • Python 3.11+

AEH OpenShell: set up a new namespace

Assume Forge SAW is already deployed by the user in [NAMESPACE]. Do not redeploy the platform or use another namespace's credentials. Follow the complete namespace setup and trigger guide before creating a PipelineRun:

  1. Install the matching OpenShell Pipeline, Tasks and pipeline ServiceAccount/RBAC in [NAMESPACE]; choose explicit Flow and OpenShell-capable harness revisions.
  2. Reuse Forge's native mTLS gateway on TCP 17670, matching client certificates and upstream CA. Map a certificate-valid hostname to the discovered Service IP in CI pods. Keep verification enabled; no external Route or TLS bridge is needed.
  3. Configure namespace-local ingress/egress and EgressFirewall allows for the gateway, DNS, GitHub/releases, Python packages, model endpoints, results services and Kubernetes API cleanup. Preserve existing Forge rules and persist additions.
  4. Provision/configure LiteLLM, MLflow, PostgreSQL and MinIO plus their Secrets. Apply results DB migrations; verify MinIO S3 access on 9000 and MLflow artifacts.
  5. Validate a temporary sandbox's skills and a real GLM response, then trigger abevalflow-pipeline-openshell and confirm evaluate and store succeed.

The validated template uses an explicit GHCR image digest, not latest. Use the image/profile/providers validated for your Forge deployment. Cases and judges come from submissions/openclaw-forge/eval.yaml and its cases/ directory in the chosen Flow revision, not an automatically executed harness bootstrap script. This profile skips the test cube; MLflow publication occurs within evaluate, while store writes evaluation_runs and MinIO artifacts. A green TCP probe alone does not validate the agent, judging or storage.

Documentation

License

Apache License 2.0

About

Automated Tekton pipeline on OpenShift for evaluating AI artifacts submissions

Resources

Stars

8 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages