You need: LLM-specific tracing (spans per model call with token counts, latency, model version) via Langfuse, LangSmith, or Arize Phoenix; standard infra monitoring (CPU/GPU, memory, API error rates); cost dashboards; and a way to log and search production prompts and completions for debugging. OpenTelemetry is increasingly standard for the trace layer.
Back to All Posts
What observability tooling do I need for AI applications?
Trusted by enterprises building the future
Add Comment