Spaces:
Sleeping
Sleeping
|
Download README.md from ibm-research/BPO-Bench: direct link, hf CLI and curl.
- Browser
- Download file 7.62 kB
-
https://huggingface.co/spaces/ibm-research/BPO-Bench/resolve/main/README.md
- Command line
-
hf download hf://spaces/ibm-research/BPO-Bench/README.md
-
curl -L -o README.md https://huggingface.co/spaces/ibm-research/BPO-Bench/resolve/main/README.md
7.62 kB
metadata
title: BPO Benchmark Evaluation
emoji: π
colorFrom: blue
colorTo: purple
sdk: docker
app_port: 7860
pinned: false
license: apache-2.0
BPO Benchmark Evaluation
Evaluate CUGA SDK on BPO (Business Process Outsourcing) recruiting analytics tasks.
What This Benchmarks
This Space evaluates the CUGA SDK - a production AI agent framework - on its ability to use tool APIs to answer recruiting analytics questions.
Features
- Real Agent Testing: Uses the actual CUGA SDK from PyPI (
pip install cuga>=0.2.8) - 32 Tool APIs for recruiting data analysis (13 core + 19 error-prone)
- 45 Evaluation Tasks across 6 test suites (easy/medium/hard difficulty)
- Multi-Metric Scoring: String similarity, exact match, keyword match, API accuracy, and LLM Judge
- OpenAI and Groq provider support
- Langfuse Observability: Optional tracing for detailed evaluation analysis
API Endpoints
The benchmark exposes 32 BPO recruiting analytics endpoints plus a health check.
Core Endpoints (13)
Candidate Source APIs (7)
- SLA performance by source
- Total hires by source
- Candidate volume metrics
- Funnel conversion rates
- Metadata and timeframe
- Definitions and methodology
- Source recommendations
Skills APIs (6)
- Skill analysis and correlations
- Skill impact on fill rate
- Skill impact on SLA
- Skill relevance justification
- Success criteria
- Data sources used
Error-Prone Endpoints (19)
These endpoints intentionally exhibit problematic behaviors to test agent resilience.
Type Mismatch (3)
skills/skill-summaryβ Returns plain string instead of JSON objectcandidate-source/source-sla-scoreβ Returns numeric float instead of structured responsecandidate-source/inactive-sourcesβ Returns boolean or list depending on data state
HTTP Errors (4)
candidate-source/candidate-pipeline-statusβ Intermittently returns 404candidate-source/source-sla-checkβ Returns 500 Internal Server Errorcandidate-source/funnel-statusβ Returns 503 Service Unavailablecandidate-source/bulk-source-dataβ Returns 429 Too Many Requests
Schema Violations (4)
skills/model-registryβ No Pydantic output schema; returns untyped dictskills/skill-lookupβ Returns extra undeclared fieldscandidate-source/source-metrics-liteβ Randomly omits required fieldscandidate-source/volume-reportβ Returns wrong field types (strings for numbers)
Edge Cases (5)
candidate-source/full-candidate-detailsβ Returns oversized payload (~1MB)candidate-source/source-directoryβ Contains Unicode and special charactersskills/skill-deep-analysisβ Deeply nested JSON (5+ levels)candidate-source/sla-extendedβ Includes unexpected extra fieldsskills/analyze-skill-matchβ Returns mismatched schema vs documentation
Undocumented Behaviors (3)
candidate-source/requisition-detailsβ Non-standard error format (error as string, not object)candidate-source/list-all-sourcesβ Undocumented pagination in responsecandidate-source/batch-metricsβ Undocumented rate limiting headers
Usage
- Enter your OpenAI or Groq API key
- Select test suites to evaluate (checkboxes):
- Core (26 tasks) β Standard recruiting analytics questions
- Type Mismatch (3 tasks) β APIs returning unexpected data types
- HTTP Errors (4 tasks) β APIs returning HTTP error codes
- Schema Violations (4 tasks) β APIs with missing/wrong schema fields
- Edge Cases (5 tasks) β Large payloads, Unicode, deep nesting
- Undocumented Behaviors (3 tasks) β Non-standard error formats and pagination
- Optionally filter to specific task IDs within selected suites
- Click Run Evaluation
- View results with multi-metric scoring and per-task breakdowns
Evaluation Metrics
Each task is scored on multiple dimensions:
- String Similarity: Fuzzy match between agent response and expected output (0-100%)
- Exact Match: Binary check for precise answer correctness
- Keyword Match Rate: Percentage of expected keywords found in the response
- API Accuracy: Precision and recall of which API endpoints the agent invoked
- LLM Judge (optional): GPT-based semantic evaluation of response quality
- Final Composite Score: Weighted combination of all metrics
A task passes when the composite score exceeds the threshold.
Dataset
The benchmark uses synthetic BPO recruiting data:
- 64k candidate records
- 1,047 requisitions
- 7 sourcing channels
Dataset: ibm-research/BPO-Bench
Local Development
# Clone the space
git clone https://huggingface.co/spaces/ibm-research/BPO-Bench
cd BPO-Bench
# Install dependencies
pip install -r requirements.txt
# Download data (all task suites + fixture data)
python -c "
from huggingface_hub import hf_hub_download
import os
os.makedirs('data', exist_ok=True)
files = [
'candidate_data.parquet',
'tasks.json',
'tasks_type_mismatch.json',
'tasks_http_errors.json',
'tasks_schema_violations.json',
'tasks_edge_cases.json',
'tasks_undocumented.json',
'large_response_fixture.json',
]
for f in files:
hf_hub_download('ibm-research/BPO-Bench', f, local_dir='data', repo_type='dataset')
print(f'Downloaded {len(files)} files')
"
# Run
python app.py
Architecture
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Gradio UI (port 7860) β
β - 6 test suite checkboxes, multi-metric result display β
β - LLM Judge toggle, Langfuse observability β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β CUGA SDK Agent β
β - Loads tools via OpenAPI spec from Registry β
β - Processes queries with LLM (OpenAI/Groq) β
β - Orchestrates tool calls β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βββββββββββ΄βββββββββββ
βΌ βΌ
ββββββββββββββββββββββββββββ ββββββββββββββββββββββββββββ
β FastAPI Server (:8000) β β CUGA Registry (:8001) β
β - 32 BPO API endpoints β β - Tool discovery β
β - OpenAPI spec β β - OpenAPI aggregation β
β - Loads data from β ββββββββββββββββββββββββββββ
β parquet β
ββββββββββββββββββββββββββββ
License
Apache 2.0