Calibrated Multi-Task Tool Classifier (DeBERTa-v3-Large Tri-Axis)

A fine-tuned multi-task microsoft/deberta-v3-large foundation model (434,018,310 parameters) for ahead-of-time semantic classification of tool schemas in autonomous agent runtimes.

Part of the Agent Runtime Optimization Middleware architecture.

Architecture & Multi-Task Heads

Rather than deploying multiple separate models, this architecture shares a single DeBERTa-v3-Large encoder backbone and projects into three task-specific classification heads (Linear(1024 -> 2) each, 6,150 head parameters total, 99.9986% shared backbone):

  1. State Mutability (effect):
    • READ (0): Execution induces no persistent side-effects (cacheable).
    • WRITE (1): Execution mutates persistent state (enforces sequential write barrier).
  2. Execution Boundary (medium):
    • LOCAL_INTERNAL (0): Process/sandbox local (long-horizon cache).
    • EXTERNAL_API (1): Remote network/SaaS API (bounded TTL + revalidation).
  3. Temporal Invariance (volatility):
    • DETERMINISTIC (0): Output invariant across identical arguments.
    • VOLATILE (1): Non-stationary / time-dependent / stochastic (never cached, TTL=0.0).

Latency & Serving Characteristics (Measured)

  • Single-Call Serving Latency (Batch 1): 23.9โ€“32.7 ms (dynamic padding: 20.4โ€“27.9 ms).
  • Batched Prewarm Throughput: 7.5โ€“9.7 ms / tool (at batch size 32).
  • Evaluation Throughput (Batch 16): 11.0 ms / tool.
  • Weights Checkpoint: 1.74 GB (tri_axis_model.pt, fp32).
  • In-Memory / SQLite Registry Hit: 0.0002 ms (0 DNN invocations on hot paths after memoization).

Empirical Performance (Test Split: 406 Tools)

All metrics independently reproduced and verified against checkpoint weights:

Axis Accuracy True Macro F1 Binary F1 (Pos. Class) Platt Temp ($T$) Calibrated ECE
State Mutability (Effect) 98.03% 0.9790 0.9739 1.5035 0.0145
Boundary (Medium) 98.52% 0.9833 0.9890 1.3567 0.0129
Temporal Invariance (Volatility) 99.26% 0.9926 0.9926 1.3429 0.0052

Note on Macro F1: Reported values represent true unweighted Macro F1 $((F_{1, \text{class } 0} + F_{1, \text{class } 1}) / 2)$. Positive-class binary F1 is also provided for clarity.

Calibration & Safe-Fail Rule

  • Platt Temperature Calibration: Fitted via L-BFGS minimizing cross-entropy validation loss ($T_{\text{eff}}=1.5035, T_{\text{med}}=1.3567, T_{\text{vol}}=1.3429$).
  • Asymmetric Safe-Fail Barrier: $$\text{Effect} = \begin{cases} \text{READ} & \text{if } P(\text{READ}) \ge 0.95 \ \text{WRITE} & \text{otherwise} \end{cases}$$ Borderline or uncertain predictions default to WRITE to prevent cache-induced state corruption.

Training Corpora & Data Provenance

  • Dataset Size: 4,000 multi-domain API tool schemas (3,197 train, 397 validation, 406 test).
  • Data Sources: ToolBench, Berkeley Function-Calling Leaderboard (BFCL), SEAL-Tools, Salesforce APIGen (xLAM), and Glaive-v2.
  • Provenance Performance Split:
    • Synthesized Tool Schemas (305 test rows): 100.0% accuracy across all three heads (0 errors).
    • Real-World Natural Tools (101 test rows): 92.1% Effect, 94.1% Medium, 97.0% Volatility.
  • Volatility Domain Shift & Safeguards: On natural tool APIs without explicit entropy or polling tokens in docstrings, the middleware couples this classifier with Tier-0 AST inspection and keyword fail-safe overrides to ensure non-deterministic tools are never cached.

Model Usage

from transformers import AutoTokenizer
import torch

REPO_ID = "DamonRicci/agent-tool-tri-axis-deberta-v3-large"
tokenizer = AutoTokenizer.from_pretrained(REPO_ID)

# In the Agent Runtime Middleware, use TriAxisToolClassifier:
from agent_runtime_middleware.classifier import TriAxisToolClassifier

classifier = TriAxisToolClassifier(model_dir="./models/deberta_v3_large_tri_axis")
result = classifier.classify_tool(
    tool_name="get_order_status",
    docstring="Fetches the current processing status of an order.",
    parameters=["order_id"],
    domain="E-Commerce"
)

print(result.effect)      # READ or WRITE
print(result.medium)      # LOCAL_INTERNAL or EXTERNAL_API
print(result.volatility)  # DETERMINISTIC or VOLATILE
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support