← Back to all spotlights

Langfuse: Open-Source LLM Observability, Traces, and Evals

Debug agentic loops, audit token spend, and version prompts locally with Langfuse, the open-source engineering platform built for production AI.

P24
By Pickwise24 Editorial Team
Verified Open-Source Review

What Is Langfuse?

Langfuse is an open-source LLM engineering and observability platform that provides production tracing, prompt management, cost tracking, and automated evaluation for applications powered by Large Language Models.

>

* GitHub Repository: langfuse/langfuse

* Primary Architecture: TypeScript / Next.js frontend, PostgreSQL (metadata and configurations), ClickHouse (high-throughput trace analytics), and OpenTelemetry-compliant ingestion engines.

* Licence: Open Source (MIT / Fair Source)


Every developer building with large language models shares the same recurring nightmare: deploying an autonomous agent on Friday afternoon, only to wake up on Saturday morning to discover it spent the night stuck in an existential infinite loop, querying Claude 3.5 Sonnet four hundred times to determine whether a user's date format is ISO-compliant.

Traditional application performance monitoring (APM) tools like Datadog or Sentry excel at telling you that an HTTP endpoint returned a 500 Internal Server Error. What they cannot tell you is why your Retrieval-Augmented Generation (RAG) pipeline hallucinated an apology, why latency spiked by four seconds on step three of an agentic tool chain, or which specific prompt template revision broke production output.

Langfuse solves this visibility problem cleanly without forcing you into proprietary SaaS vendor lock-in.


+--------------------------------------------------------------------+
|                         Langfuse Platform                          |
+--------------------------------------------------------------------+
|  [Ingestion API]  <-- OpenTelemetry / Python / TS SDKs / LangChain |
|         |                                                          |
|         +--> [PostgreSQL]  (Prompts, Users, Config, Eval Scores)   |
|         |                                                          |
|         +--> [ClickHouse]  (Traces, Spans, Generations, Token Logs)|
+--------------------------------------------------------------------+
|                       Web UI & Debugger                            |
|    - Visual Trace Trees   - Latency/Cost Graphs   - Prompt Sandbox |
+--------------------------------------------------------------------+

Architectural Breakdown: Built for High-Volume Ingestion

The biggest complaint across developer forums regarding self-hosted telemetry tools is overhead. If your observability layer adds 120 milliseconds of latency to an already sluggish generation call, it has failed its primary duty.

Langfuse sidesteps this problem through asynchronous batching and a dual-storage paradigm:

1. Transactional Integrity (PostgreSQL): Prompt versions, role-based access control, organisations, and evaluation criteria live in relational Postgres. If you update a production prompt template, the operation is ACID-compliant and instant.

2. Analytical Throughput (ClickHouse): When processing millions of trace events, vector distances, and raw token counts, standard relational databases buckle. Langfuse uses ClickHouse under the bonnet for analytical queries, enabling lightning-fast trace filtering across billions of tokens without degrading API ingest performance.

3. Open Standards: Built natively around the OpenTelemetry (OTel) specification, Langfuse doesn't invent its own bizarre tracing taxonomy. A "Trace" represents the full request cycle, composed of nested "Spans" (internal code blocks) and "Generations" (specific model calls with input, output, and token accounting).


Observability Tool Comparison

FeatureLangfuseLangSmithArize Phoenix
Licence ModelOpen Source (MIT core)Closed Source (SaaS-first)Open Source (Elastic License)
Self-HostingStraightforward (docker-compose)Complex enterprise deployLightweight local / container
Prompt ManagementBuilt-in with SDK fetch & cacheBuilt-inLimited
Telemetry StandardNative OpenTelemetryProprietary RunTreeNative OpenTelemetry
Storage BackendPostgreSQL + ClickHouseProprietary CloudIn-memory / DuckDB / Postgres

Local Setup: Up and Running in 60 Seconds

You can spin up a fully featured, self-hosted Langfuse stack locally using the official Docker Compose manifest.


# Clone the repository
git clone https://github.com/langfuse/langfuse.git
cd langfuse

# Start the complete stack (Langfuse Web, Worker, PostgreSQL, ClickHouse, MinIO)
docker compose up -d

Navigate to http://localhost:3000, register an admin account, create a new project, and retrieve your API keys (pk-lf-... and sk-lf-...).


Instrumenting Your Code

Integrating Langfuse into an existing Python pipeline requires minimal refactoring. You can either use their drop-in wrapper around OpenAI/Anthropic clients, or utilise their native decorators.

Python Example: Automatic Tracing & Cost Auditing


import os
from langfuse.openai import openai
from langfuse.decorators import observe

# Environment variables
os.environ["LANGFUSE_PUBLIC_KEY"] = "pk-lf-..."
os.environ["LANGFUSE_SECRET_KEY"] = "sk-lf-..."
os.environ["LANGFUSE_HOST"] = "http://localhost:3000"

@observe()
def retrieve_context(query: str) -> str:
    # Simulate a vector database retrieval step
    return "Langfuse provides open-source tracing and prompt management."

@observe()
def generate_response(user_query: str) -> str:
    context = retrieve_context(user_query)
    
    # Langfuse wraps the OpenAI client to capture inputs, outputs, 
    # tokens, and exact cost calculations transparently
    response = openai.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {"role": "system", "content": f"Answer using context: {context}"},
            {"role": "user", "content": user_query}
        ]
    )
    return response.choices[0].message.content

if __name__ == "__main__":
    result = generate_response("What does Langfuse do?")
    print(f"Agent Output: {result}")

Every invocation automatically dispatches an asynchronous trace in the background. If you inspect the web UI, you will see a detailed visual breakdown detailing the exact millisecond duration of the retrieve_context span, the prompt sent to gpt-4o-mini, and the precise fractional cent spent on the request.


Core Feature Highlights

1. In-Code Prompt Management with Dynamic Fallbacks

Hardcoding prompts in application repositories leads to endless redeployments simply to tweak a single instruction. Langfuse lets you manage prompts via their UI while maintaining local caching in your application:


from langfuse import Langfuse

langfuse = Langfuse()

# Fetches prompt from Langfuse with local in-memory fallback
prompt = langfuse.get_prompt("summariser_system_prompt", label="production")
compiled_template = prompt.compile(max_words=50)

If your telemetry server ever becomes unreachable, the client falls back to cached versions rather than terminating user requests.

2. Multi-Turn Agentic Tracing

Modern architectures frequently chain multiple agents, tools, and retrievers together. Langfuse flattens nested, recursive calls into an intuitive Gantt-style tree view, pinpointing exactly which tool stalled an execution flow.

3. Systematic Evals (LLM-as-a-Judge)

You can establish programmatic evaluations that run automatically on production samples. Evaluate responses for hallucination, brand tone, or toxicity using off-the-shelf judge models, or attach human feedback scores via the UI to curate clean fine-tuning datasets.


Key Takeaways

  • Data Sovereignty: Langfuse is entirely self-hostable, making it a viable option for engineering teams under strict GDPR, HIPAA, or SOC2 data-residency requirements.
  • Cost Efficiency: It tracks exact token prices across hundreds of models out of the box, preventing billing surprises in autonomous loops.
  • Standardised Ingestion: Built on OpenTelemetry, removing lock-in risks associated with proprietary SDKs.
  • Production-Grade Analytics: The ClickHouse-backed trace engine handles real-world traffic volumes without bogging down query rendering.

🛡️ Editorial Standards & Methodology

Every repository featured on Pickwise24 undergoes testing on local workstation hardware before publication. We verify CLI installation steps, review open-source repository licensing, benchmark computational footprint, and evaluate architectural trade-offs to provide genuine, high-utility developer intelligence.