Most engineering teams approach LLM integration with an initial burst of optimism, followed rapidly by architectural regret. What starts as a tidy Python wrapper around an API endpoint invariably degrades into a 2,000-line monolith juggling retries, rate limits, schema validations, prompt templates, and cost telemetry. If one upstream provider has a hiccup, the user interface hangs, and debugging involves combing through unstructured application logs like a digital archaeologist.
Enter TensorZero (github.com/tensorzero/tensorzero), an open-source, high-performance LLM gateway written in Rust designed to sit between your client applications and model providers. Rather than treating model interactions as unpredictable external webhooks, TensorZero treats inference as an optimisable feedback loop: parsing requests, managing dynamic fallback routing, recording structured feedback, and packaging that data directly for downstream model evaluation and fine-tuning.
┌──────────────────────────────────────────────────────────┐
│ Application Layer │
└────────────────────────────┬─────────────────────────────┘
│ (Clean Function Call)
▼
┌──────────────────────────────────────────────────────────┐
│ TensorZero Gateway │
│ ┌────────────────────────────────────────────────────┐ │
│ │ Rust Core: Schemas • Caching • Fallback Routing │ │
│ └─────────────────────────┬──────────────────────────┘ │
└────────────┬───────────────┴───────────────┬─────────────┘
│ │
▼ ▼
┌───────────────────────────┐ ┌──────────────────────────┐
│ External Model APIs │ │ ClickHouse Database │
│ (Anthropic, OpenAI, etc.) │ │ (Inference & Feedback) │
└───────────────────────────┘ └────────────┬─────────────┘
│
▼
┌──────────────────────────┐
│ Evaluation & Fine-Tuning │
└──────────────────────────┘
What Is TensorZero?
TensorZero is an open-source LLM gateway and data-flywheel engine engineered in Rust. It abstracts model providers behind declarative configuration files (tensorzero.toml), manages structured JSON inputs and outputs with strict validation, and logs every interaction directly into ClickHouse to fuel continuous automated evaluation and supervised fine-tuning.
Key Takeaways
- Language-Agnostic Core: Written in Rust, it delivers microsecond-level gateway overhead alongside native multi-threading.
- Declarative Function Routing: Prompts, fallback chains, and schema constraints live in version-controlled config files, not scattered through application code.
- Continuous Learning Loop: Binds user feedback and output metrics directly to raw inference logs inside an embedded or external ClickHouse instance.
- Provider Redundancy: Automatically degrades to backup models or alternative providers upon timeout, rate-limiting, or schema validation failures.
Architectural Deep Dive: Moving Prompts Out of the Codebase
In standard AI applications, changing a prompt or swapping gpt-4o for claude-3-5-sonnet often demands a pull request, continuous integration runs, and a full application redeploy. TensorZero separates the business logic (what information your code needs) from the model variant (how that information is extracted).
Your code executes an abstract function (for example, generate_product_summary). The TensorZero gateway resolves that function against a declarative configuration, applies the relevant system prompts, validates the response against strict JSON Schemas, and routes the request across your configured providers.
TensorZero vs Generic Reverse Proxies
| Feature | Generic API Proxy (LiteLLM, Kong) | TensorZero |
|---|---|---|
| Core Language | Python / Lua / Go | Rust |
| Data Engine | Postgres / Redis / Stateless | Embedded or Managed ClickHouse |
| Prompt Versioning | Application-level or manual dashboard | Declarative TOML configuration |
| Fine-Tuning Loop | Export scripts required | Native feedback collection & dataset curation |
| Schema Validation | Best-effort string matching | Compile-time/Runtime JSON Schema enforcement |
Hands-On: Setting Up TensorZero Locally
Getting a local instance running takes under three minutes using Docker Compose, which pulls the compiled Rust gateway binary alongside a local ClickHouse instance.
1. Initialise the Project Directory
mkdir tensorzero-quickstart && cd tensorzero-quickstart
curl -sSL https://raw.githubusercontent.com/tensorzero/tensorzero/main/examples/quickstart/docker-compose.yml -o docker-compose.yml
curl -sSL https://raw.githubusercontent.com/tensorzero/tensorzero/main/examples/quickstart/tensorzero.toml -o tensorzero.toml
2. Define Your Function in tensorzero.toml
Here we define a structured function that extracts actionable tasks from meeting transcripts, with explicit fallbacks across providers.
[functions.extract_tasks]
type = "json"
system_template = "You extract actionable tasks from transcripts. Always return valid JSON."
[functions.extract_tasks.output_schema]
type = "object"
required = ["tasks"]
properties = { tasks = { type = "array", items = { type = "string" } } }
# Primary model variant
[functions.extract_tasks.variants.claude_primary]
model = "anthropic::claude-3-5-sonnet-20241022"
# Fallback model variant if rate-limited or unavailable
[functions.extract_tasks.variants.openai_fallback]
model = "openai::gpt-4o-mini"
3. Spin Up the Gateway
Supply your provider keys and boot the stack:
export ANTHROPIC_API_KEY="your-anthropic-key"
export OPENAI_API_KEY="your-openai-key"
docker compose up -d
Executing Inferences and Recording Feedback
Once the gateway is listening on port 3000, client libraries (available for Python, TypeScript, or raw HTTP) communicate with TensorZero rather than hitting model APIs directly.
Running Inference via cURL
curl -X POST http://localhost:3000/inference \
-H "Content-Type: application/json" \
-d '{
"function_name": "extract_tasks",
"input": {
"messages": [
{
"role": "user",
"content": "Alice: I will fix the staging bug. Bob: Brilliant, I will draft the release notes."
}
]
}
}'
The gateway returns a validated payload containing the model's structured output alongside an inference_id:
{
"inference_id": "018f2940-1a67-7890-bcde-abcdef123456",
"output": {
"tasks": [
"Alice to fix the staging bug",
"Bob to draft the release notes"
]
},
"variant_name": "claude_primary"
}
The Flywheel: Closing the Feedback Loop
The true value proposition of TensorZero surfaces when user actions or downstream test runners evaluate the result. You can submit real-time evaluations back to the gateway using that inference_id:
curl -X POST http://localhost:3000/feedback \
-H "Content-Type: application/json" \
-d '{
"inference_id": "018f2940-1a67-7890-bcde-abcdef123456",
"metric_name": "task_extraction_accuracy",
"value": 1.0
}'
Because inferences and feedback write directly into ClickHouse tables, you bypass the painful step of manually building data pipelines to assemble training pairs. When fine-tuning a small open-weights model (such as Llama 3 or Mistral) to replace an expensive frontier model, TensorZero lets you filter high-scoring historical calls directly from your storage layer to build your training splits.
Why TensorZero Deserves Your Attention
The broader AI ecosystem has spent the past eighteen months obsessing over toy agent frameworks while quietly sweeping reliability engineering under the rug. In high-throughput environments, Python-based reverse proxies frequently become memory hogs under sustained concurrent connections, introducing jitter that poisons time-to-first-token benchmarks.
By betting on Rust and ClickHouse, TensorZero removes the latency penalty of running an intermediary proxy. It abstracts away the vendor lock-in that makes pivoting between Anthropic, OpenAI, or self-hosted vLLM instances such a chore, while laying the data foundation required to systematically reduce token costs over time through fine-tuning.