--- title: BPO Benchmark Evaluation emoji: "\U0001F4CA" colorFrom: blue colorTo: purple sdk: docker app_port: 7860 pinned: false license: apache-2.0 --- # BPO Benchmark Evaluation Evaluate **CUGA SDK** on BPO (Business Process Outsourcing) recruiting analytics tasks. ## What This Benchmarks This Space evaluates the [CUGA SDK](https://pypi.org/project/cuga/) - a production AI agent framework - on its ability to use tool APIs to answer recruiting analytics questions. ## Features - **Real Agent Testing**: Uses the actual CUGA SDK from PyPI (`pip install cuga>=0.2.8`) - **32 Tool APIs** for recruiting data analysis (13 core + 19 error-prone) - **45 Evaluation Tasks** across 6 test suites (easy/medium/hard difficulty) - **Multi-Metric Scoring**: String similarity, exact match, keyword match, API accuracy, and LLM Judge - **OpenAI and Groq** provider support - **Langfuse Observability**: Optional tracing for detailed evaluation analysis ## API Endpoints The benchmark exposes 32 BPO recruiting analytics endpoints plus a health check. ### Core Endpoints (13) #### Candidate Source APIs (7) - SLA performance by source - Total hires by source - Candidate volume metrics - Funnel conversion rates - Metadata and timeframe - Definitions and methodology - Source recommendations #### Skills APIs (6) - Skill analysis and correlations - Skill impact on fill rate - Skill impact on SLA - Skill relevance justification - Success criteria - Data sources used ### Error-Prone Endpoints (19) These endpoints intentionally exhibit problematic behaviors to test agent resilience. #### Type Mismatch (3) - `skills/skill-summary` — Returns plain string instead of JSON object - `candidate-source/source-sla-score` — Returns numeric float instead of structured response - `candidate-source/inactive-sources` — Returns boolean or list depending on data state #### HTTP Errors (4) - `candidate-source/candidate-pipeline-status` — Intermittently returns 404 - `candidate-source/source-sla-check` — Returns 500 Internal Server Error - `candidate-source/funnel-status` — Returns 503 Service Unavailable - `candidate-source/bulk-source-data` — Returns 429 Too Many Requests #### Schema Violations (4) - `skills/model-registry` — No Pydantic output schema; returns untyped dict - `skills/skill-lookup` — Returns extra undeclared fields - `candidate-source/source-metrics-lite` — Randomly omits required fields - `candidate-source/volume-report` — Returns wrong field types (strings for numbers) #### Edge Cases (5) - `candidate-source/full-candidate-details` — Returns oversized payload (~1MB) - `candidate-source/source-directory` — Contains Unicode and special characters - `skills/skill-deep-analysis` — Deeply nested JSON (5+ levels) - `candidate-source/sla-extended` — Includes unexpected extra fields - `skills/analyze-skill-match` — Returns mismatched schema vs documentation #### Undocumented Behaviors (3) - `candidate-source/requisition-details` — Non-standard error format (error as string, not object) - `candidate-source/list-all-sources` — Undocumented pagination in response - `candidate-source/batch-metrics` — Undocumented rate limiting headers ## Usage 1. Enter your **OpenAI** or **Groq** API key 2. Select test suites to evaluate (checkboxes): - **Core** (26 tasks) — Standard recruiting analytics questions - **Type Mismatch** (3 tasks) — APIs returning unexpected data types - **HTTP Errors** (4 tasks) — APIs returning HTTP error codes - **Schema Violations** (4 tasks) — APIs with missing/wrong schema fields - **Edge Cases** (5 tasks) — Large payloads, Unicode, deep nesting - **Undocumented Behaviors** (3 tasks) — Non-standard error formats and pagination 3. Optionally filter to specific task IDs within selected suites 4. Click **Run Evaluation** 5. View results with multi-metric scoring and per-task breakdowns ## Evaluation Metrics Each task is scored on multiple dimensions: - **String Similarity**: Fuzzy match between agent response and expected output (0-100%) - **Exact Match**: Binary check for precise answer correctness - **Keyword Match Rate**: Percentage of expected keywords found in the response - **API Accuracy**: Precision and recall of which API endpoints the agent invoked - **LLM Judge** (optional): GPT-based semantic evaluation of response quality - **Final Composite Score**: Weighted combination of all metrics A task **passes** when the composite score exceeds the threshold. ## Dataset The benchmark uses synthetic BPO recruiting data: - 64k candidate records - 1,047 requisitions - 7 sourcing channels **Dataset:** [ibm-research/BPO-Bench](https://huggingface.co/datasets/ibm-research/BPO-Bench) ## Local Development ```bash # Clone the space git clone https://huggingface.co/spaces/ibm-research/BPO-Bench cd BPO-Bench # Install dependencies pip install -r requirements.txt # Download data (all task suites + fixture data) python -c " from huggingface_hub import hf_hub_download import os os.makedirs('data', exist_ok=True) files = [ 'candidate_data.parquet', 'tasks.json', 'tasks_type_mismatch.json', 'tasks_http_errors.json', 'tasks_schema_violations.json', 'tasks_edge_cases.json', 'tasks_undocumented.json', 'large_response_fixture.json', ] for f in files: hf_hub_download('ibm-research/BPO-Bench', f, local_dir='data', repo_type='dataset') print(f'Downloaded {len(files)} files') " # Run python app.py ``` ## Architecture ``` ┌─────────────────────────────────────────────────────────────┐ │ Gradio UI (port 7860) │ │ - 6 test suite checkboxes, multi-metric result display │ │ - LLM Judge toggle, Langfuse observability │ └─────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────┐ │ CUGA SDK Agent │ │ - Loads tools via OpenAPI spec from Registry │ │ - Processes queries with LLM (OpenAI/Groq) │ │ - Orchestrates tool calls │ └─────────────────────────────────────────────────────────────┘ │ ┌─────────┴──────────┐ ▼ ▼ ┌──────────────────────────┐ ┌──────────────────────────┐ │ FastAPI Server (:8000) │ │ CUGA Registry (:8001) │ │ - 32 BPO API endpoints │ │ - Tool discovery │ │ - OpenAPI spec │ │ - OpenAPI aggregation │ │ - Loads data from │ └──────────────────────────┘ │ parquet │ └──────────────────────────┘ ``` ## License Apache 2.0