Skip to content

Trace agents with Langfuse

Level: Application

Outcome

Trace an Anvia agent in Langfuse, attach stable product context, flush short-lived work, and publish evaluation scores against the originating trace.

When to use it

Use @anvia/langfuse when Langfuse is your tracing, prompt, dataset, or scoring platform. The adapter is server-side; never include Langfuse secret keys in a browser bundle.

Flow

text
Anvia observer -> @anvia/langfuse -> Langfuse trace/generation/tool observations
Anvia eval reporter ----------------> trace-correlated score

Setup

sh
pnpm add @anvia/core @anvia/openai @anvia/langfuse

Set LANGFUSE_PUBLIC_KEY, LANGFUSE_SECRET_KEY, LANGFUSE_BASE_URL, and optionally environment, release, and service-name variables in server secrets.

Process and request boundaries

ts
import { AgentBuilder } from "@anvia/core/agent";
import { langfuse } from "@anvia/langfuse";
import { OpenAIClient } from "@anvia/openai";

const tracing = langfuse.create({
  serviceName: "support-api",
  captureMode: "safe",
});

const openai = new OpenAIClient({ apiKey: process.env.OPENAI_API_KEY });
const agent = new AgentBuilder("support", openai.completionModel("gpt-5"))
  .observe(tracing)
  .instructions("Answer with verified support policy only.")
  .build();

try {
  const response = await agent
    .prompt("Who may change billing settings?")
    .withTrace({
      name: "support-answer",
      userId: "usr_opaque_42",
      sessionId: "conv_opaque_91",
      tags: ["support"],
      metadata: { channel: "web" },
    })
    .send();

  await tracing.flush(); // useful for a short-lived command or job
  console.log(response.output, response.trace?.traceId);
} finally {
  await tracing.shutdown();
}

For a long-running server, create one adapter per process and flush through normal batching. Shut it down once during graceful process termination, not after every request.

Expected behavior and failures

Langfuse receives a root run plus model and tool observations. send() completing does not guarantee that buffered telemetry is already delivered. Configuration may be valid while the network or keys fail only during export, flush, or shutdown.

Privacy, security, and production adaptations

Safe capture still exports identifiers, tags, metadata, usage, timing, and exception information. Use opaque IDs; never put access tokens or raw customer content in metadata. Approve retention, access, deletion, and payload capture before using full mode. Size the score queue and retry policy, monitor failed exports, and keep telemetry failure separate from agent availability unless the job requires confirmed delivery.

Tests

Mock the adapter for agent unit tests. In adapter integration tests, assert trace context, tool observation mapping, safe capture, score publishing, failed export handling, and final queue drain. Use synthetic content in a staging smoke trace.

Source and extensions

Built for Anvia.