Calibrated Multi-Task Tool Classifier (DeBERTa-v3-Large Tri-Axis)
A fine-tuned multi-task microsoft/deberta-v3-large foundation model (434,018,310 parameters) for ahead-of-time semantic classification of tool schemas in autonomous agent runtimes.
Part of the Agent Runtime Optimization Middleware architecture.
Architecture & Multi-Task Heads
Rather than deploying multiple separate models, this architecture shares a single DeBERTa-v3-Large encoder backbone and projects into three task-specific classification heads (Linear(1024 -> 2) each, 6,150 head parameters total, 99.9986% shared backbone):
- State Mutability (
effect):READ(0): Execution induces no persistent side-effects (cacheable).WRITE(1): Execution mutates persistent state (enforces sequential write barrier).
- Execution Boundary (
medium):LOCAL_INTERNAL(0): Process/sandbox local (long-horizon cache).EXTERNAL_API(1): Remote network/SaaS API (bounded TTL + revalidation).
- Temporal Invariance (
volatility):DETERMINISTIC(0): Output invariant across identical arguments.VOLATILE(1): Non-stationary / time-dependent / stochastic (never cached, TTL=0.0).
Latency & Serving Characteristics (Measured)
- Single-Call Serving Latency (Batch 1): 23.9โ32.7 ms (dynamic padding: 20.4โ27.9 ms).
- Batched Prewarm Throughput: 7.5โ9.7 ms / tool (at batch size 32).
- Evaluation Throughput (Batch 16): 11.0 ms / tool.
- Weights Checkpoint: 1.74 GB (
tri_axis_model.pt, fp32). - In-Memory / SQLite Registry Hit: 0.0002 ms (0 DNN invocations on hot paths after memoization).
Empirical Performance (Test Split: 406 Tools)
All metrics independently reproduced and verified against checkpoint weights:
| Axis | Accuracy | True Macro F1 | Binary F1 (Pos. Class) | Platt Temp ($T$) | Calibrated ECE |
|---|---|---|---|---|---|
| State Mutability (Effect) | 98.03% | 0.9790 | 0.9739 | 1.5035 | 0.0145 |
| Boundary (Medium) | 98.52% | 0.9833 | 0.9890 | 1.3567 | 0.0129 |
| Temporal Invariance (Volatility) | 99.26% | 0.9926 | 0.9926 | 1.3429 | 0.0052 |
Note on Macro F1: Reported values represent true unweighted Macro F1 $((F_{1, \text{class } 0} + F_{1, \text{class } 1}) / 2)$. Positive-class binary F1 is also provided for clarity.
Calibration & Safe-Fail Rule
- Platt Temperature Calibration: Fitted via L-BFGS minimizing cross-entropy validation loss ($T_{\text{eff}}=1.5035, T_{\text{med}}=1.3567, T_{\text{vol}}=1.3429$).
- Asymmetric Safe-Fail Barrier:
$$\text{Effect} = \begin{cases} \text{READ} & \text{if } P(\text{READ}) \ge 0.95 \ \text{WRITE} & \text{otherwise} \end{cases}$$
Borderline or uncertain predictions default to
WRITEto prevent cache-induced state corruption.
Training Corpora & Data Provenance
- Dataset Size: 4,000 multi-domain API tool schemas (3,197 train, 397 validation, 406 test).
- Data Sources: ToolBench, Berkeley Function-Calling Leaderboard (BFCL), SEAL-Tools, Salesforce APIGen (xLAM), and Glaive-v2.
- Provenance Performance Split:
- Synthesized Tool Schemas (305 test rows): 100.0% accuracy across all three heads (0 errors).
- Real-World Natural Tools (101 test rows): 92.1% Effect, 94.1% Medium, 97.0% Volatility.
- Volatility Domain Shift & Safeguards: On natural tool APIs without explicit entropy or polling tokens in docstrings, the middleware couples this classifier with Tier-0 AST inspection and keyword fail-safe overrides to ensure non-deterministic tools are never cached.
Model Usage
from transformers import AutoTokenizer
import torch
REPO_ID = "DamonRicci/agent-tool-tri-axis-deberta-v3-large"
tokenizer = AutoTokenizer.from_pretrained(REPO_ID)
# In the Agent Runtime Middleware, use TriAxisToolClassifier:
from agent_runtime_middleware.classifier import TriAxisToolClassifier
classifier = TriAxisToolClassifier(model_dir="./models/deberta_v3_large_tri_axis")
result = classifier.classify_tool(
tool_name="get_order_status",
docstring="Fetches the current processing status of an order.",
parameters=["order_id"],
domain="E-Commerce"
)
print(result.effect) # READ or WRITE
print(result.medium) # LOCAL_INTERNAL or EXTERNAL_API
print(result.volatility) # DETERMINISTIC or VOLATILE