BPO-Bench / README.md
haroldshipibm's picture
Upload folder using huggingface_hub
3cfa5b5 verified
|
Raw History Blame Contribute Delete
7.62 kB
metadata
title: BPO Benchmark Evaluation
emoji: πŸ“Š
colorFrom: blue
colorTo: purple
sdk: docker
app_port: 7860
pinned: false
license: apache-2.0

BPO Benchmark Evaluation

Evaluate CUGA SDK on BPO (Business Process Outsourcing) recruiting analytics tasks.

What This Benchmarks

This Space evaluates the CUGA SDK - a production AI agent framework - on its ability to use tool APIs to answer recruiting analytics questions.

Features

  • Real Agent Testing: Uses the actual CUGA SDK from PyPI (pip install cuga>=0.2.8)
  • 32 Tool APIs for recruiting data analysis (13 core + 19 error-prone)
  • 45 Evaluation Tasks across 6 test suites (easy/medium/hard difficulty)
  • Multi-Metric Scoring: String similarity, exact match, keyword match, API accuracy, and LLM Judge
  • OpenAI and Groq provider support
  • Langfuse Observability: Optional tracing for detailed evaluation analysis

API Endpoints

The benchmark exposes 32 BPO recruiting analytics endpoints plus a health check.

Core Endpoints (13)

Candidate Source APIs (7)

  • SLA performance by source
  • Total hires by source
  • Candidate volume metrics
  • Funnel conversion rates
  • Metadata and timeframe
  • Definitions and methodology
  • Source recommendations

Skills APIs (6)

  • Skill analysis and correlations
  • Skill impact on fill rate
  • Skill impact on SLA
  • Skill relevance justification
  • Success criteria
  • Data sources used

Error-Prone Endpoints (19)

These endpoints intentionally exhibit problematic behaviors to test agent resilience.

Type Mismatch (3)

  • skills/skill-summary β€” Returns plain string instead of JSON object
  • candidate-source/source-sla-score β€” Returns numeric float instead of structured response
  • candidate-source/inactive-sources β€” Returns boolean or list depending on data state

HTTP Errors (4)

  • candidate-source/candidate-pipeline-status β€” Intermittently returns 404
  • candidate-source/source-sla-check β€” Returns 500 Internal Server Error
  • candidate-source/funnel-status β€” Returns 503 Service Unavailable
  • candidate-source/bulk-source-data β€” Returns 429 Too Many Requests

Schema Violations (4)

  • skills/model-registry β€” No Pydantic output schema; returns untyped dict
  • skills/skill-lookup β€” Returns extra undeclared fields
  • candidate-source/source-metrics-lite β€” Randomly omits required fields
  • candidate-source/volume-report β€” Returns wrong field types (strings for numbers)

Edge Cases (5)

  • candidate-source/full-candidate-details β€” Returns oversized payload (~1MB)
  • candidate-source/source-directory β€” Contains Unicode and special characters
  • skills/skill-deep-analysis β€” Deeply nested JSON (5+ levels)
  • candidate-source/sla-extended β€” Includes unexpected extra fields
  • skills/analyze-skill-match β€” Returns mismatched schema vs documentation

Undocumented Behaviors (3)

  • candidate-source/requisition-details β€” Non-standard error format (error as string, not object)
  • candidate-source/list-all-sources β€” Undocumented pagination in response
  • candidate-source/batch-metrics β€” Undocumented rate limiting headers

Usage

  1. Enter your OpenAI or Groq API key
  2. Select test suites to evaluate (checkboxes):
    • Core (26 tasks) β€” Standard recruiting analytics questions
    • Type Mismatch (3 tasks) β€” APIs returning unexpected data types
    • HTTP Errors (4 tasks) β€” APIs returning HTTP error codes
    • Schema Violations (4 tasks) β€” APIs with missing/wrong schema fields
    • Edge Cases (5 tasks) β€” Large payloads, Unicode, deep nesting
    • Undocumented Behaviors (3 tasks) β€” Non-standard error formats and pagination
  3. Optionally filter to specific task IDs within selected suites
  4. Click Run Evaluation
  5. View results with multi-metric scoring and per-task breakdowns

Evaluation Metrics

Each task is scored on multiple dimensions:

  • String Similarity: Fuzzy match between agent response and expected output (0-100%)
  • Exact Match: Binary check for precise answer correctness
  • Keyword Match Rate: Percentage of expected keywords found in the response
  • API Accuracy: Precision and recall of which API endpoints the agent invoked
  • LLM Judge (optional): GPT-based semantic evaluation of response quality
  • Final Composite Score: Weighted combination of all metrics

A task passes when the composite score exceeds the threshold.

Dataset

The benchmark uses synthetic BPO recruiting data:

  • 64k candidate records
  • 1,047 requisitions
  • 7 sourcing channels

Dataset: ibm-research/BPO-Bench

Local Development

# Clone the space
git clone https://huggingface.co/spaces/ibm-research/BPO-Bench
cd BPO-Bench

# Install dependencies
pip install -r requirements.txt

# Download data (all task suites + fixture data)
python -c "
from huggingface_hub import hf_hub_download
import os
os.makedirs('data', exist_ok=True)
files = [
    'candidate_data.parquet',
    'tasks.json',
    'tasks_type_mismatch.json',
    'tasks_http_errors.json',
    'tasks_schema_violations.json',
    'tasks_edge_cases.json',
    'tasks_undocumented.json',
    'large_response_fixture.json',
]
for f in files:
    hf_hub_download('ibm-research/BPO-Bench', f, local_dir='data', repo_type='dataset')
print(f'Downloaded {len(files)} files')
"

# Run
python app.py

Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    Gradio UI (port 7860)                    β”‚
β”‚  - 6 test suite checkboxes, multi-metric result display     β”‚
β”‚  - LLM Judge toggle, Langfuse observability                 β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                              β”‚
                              β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                      CUGA SDK Agent                         β”‚
β”‚  - Loads tools via OpenAPI spec from Registry               β”‚
β”‚  - Processes queries with LLM (OpenAI/Groq)                β”‚
β”‚  - Orchestrates tool calls                                  β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                              β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β–Ό                    β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  FastAPI Server (:8000)  β”‚  β”‚  CUGA Registry (:8001)   β”‚
β”‚  - 32 BPO API endpoints  β”‚  β”‚  - Tool discovery        β”‚
β”‚  - OpenAPI spec          β”‚  β”‚  - OpenAPI aggregation    β”‚
β”‚  - Loads data from       β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚    parquet               β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

License

Apache 2.0