Field notes · 2026-09-28 · RunVouch

Langfuse vs a watchdog: you don't need tracing to know your nightly agent is broken

Tracing tells you why a run went wrong. It cannot tell you the run never happened. Where Langfuse fits, where a watchdog fits, and how to run both.

I have run agents unattended for about two years: a crypto trading bot, a handful of nightly Claude Code jobs, a weekly build that regenerates a static site. In that time I have broken things in two clearly different ways. Sometimes a run did the wrong thing, and I needed to see exactly which tool call went sideways. Sometimes a run did nothing at all, and I found out four days later because a number on a page had stopped moving.

Tracing solves the first problem well. It does not solve the second, and that is not a flaw in tracing. It is a different question.

RunVouch is a dead man's switch, cost cap and outcome check for unattended AI agents: it alerts when a scheduled run does not happen, fails, produces no evidence, loops on the same tool call, or spends more than you allowed.

What Langfuse is good at

Langfuse is an open source LLM observability platform. You instrument your application and it records the run: the prompt that went out, the response that came back, token usage, latency, retrieval steps and tool executions. Instrumentation happens inside your code, through a drop-in client wrapper, a framework callback handler, decorators, or an OpenTelemetry span processor (Langfuse tracing docs). When an agent produced a wrong answer, or a chain got slow, or one step quietly burned most of the tokens, the trace is where you find out why.

I use that kind of data. It is the reason I would not tell anyone to drop tracing. But it requires the run to have happened, and to have gotten far enough to emit spans.

Is there a Langfuse alternative for cron jobs and nightly agents?

Honestly: the question is usually framed wrong. What people mean is "I want to know my 03:00 job ran and did not cost $40", and they reach for the tool they already know.

Langfuse can get part of the way there. Monitors combine a metric, a threshold and a window such as one hour, one day or one week, and route alerts to Slack, a webhook or a GitHub Action. There is even explicit no-data handling: you can treat null as zero, keep the previous severity, show a NO_DATA state, or notify after a sustained absence of data (Langfuse monitors and alerts). So an "our traces went quiet" alarm is buildable.

Here is what that still does not give you, and why I ended up building something separate:

Tracing answers why, a watchdog answers whether

The split I use:

Put the "whether" layer outside the thing it watches, and keep it boring. Mine is a small server with SQLite behind it, and the client fails open: if the monitoring endpoint is down, the wrapped job still runs. Monitoring that can break the job is worse than no monitoring.

Lightweight LLM monitoring without running a Clickhouse cluster

If you self-host Langfuse, the documented deployment needs Postgres for transactional data, Clickhouse for traces, observations and scores, Redis or Valkey for queues and cache, and S3 or another blob store for raw events and exports, alongside the web and worker containers (Langfuse self-hosting docs). For a team with real trace volume that is a reasonable bill. For one person with six cron jobs, it is a second system to keep alive, and the failure I actually care about is a job that did not run.

RunVouch is deliberately at the other end: one FastAPI server, one SQLite file, MIT licensed if you want to self-host it, EU hosted if you do not. Free for 3 agents, $9 for Solo, $29 for Team.

What Helicone's maintenance mode says about choosing a monitoring tool

Helicone moved into maintenance mode after being acquired by Mintlify, announced in March 2026, and Langfuse now publishes a migration guide that describes it plainly: services remain live for the foreseeable future, but the product is no longer actively developed, and the advice is to export your data early (Langfuse migration guide).

Two lessons I took from that. Keep the layer that must never lie small, portable and self-hostable, because migrating it should be an afternoon and not a quarter. And be suspicious of a setup where losing your observability vendor also means losing your only signal that a nightly job stopped working.

How I wire up both on one job

Declare the agent once, then wrap the command:

rv agent nightly-report --cadence 24h --cap-run-cost 2 --evidence
rv run nightly-report --evidence-file out/report.html -- claude -p "$(cat brief.md)"

The cadence is the dead man's switch: no check-in inside 24 hours and I get a MISSED alert on Telegram, Slack or a webhook. The evidence file is the outcome check: exit code zero but an untouched output file is NO_EVIDENCE, not success. The cap is the brake: crossing it pauses the agent, and the next rv run is refused. It never kills a running process, because killing a half-finished trading or deploy step is its own kind of damage.

Inside Claude Code I add the plugin, which hangs off the SessionStart, PostToolUse and Stop hooks (Claude Code hooks reference) and reads tokens and cost from the transcript. There is an MCP server in the official registry as com.runvouch/runvouch, plus Python and Node clients. For anything that cannot host a client, an n8n branch, a Home Assistant automation, a shell one-liner on a NAS, each agent has a ping URL: /ping/<token> to report success, /ping/<token>/start, /ping/<token>/fail, or the exit code appended. Anything that can call a URL can check in.

The detectors that run on top of those check-ins: MISSED, FAILED, NO_EVIDENCE, RETRY_STORM (the same tool with identical input 8 times or more), BUDGET_RUN, BUDGET_DAY, DRIFT (runtime or cost far off the last 7 runs, by median absolute deviation) and STALLED. When one of those fires, that is when I open the trace.

FAQ

Should I replace Langfuse with a watchdog? No. If you are debugging agent quality, prompt regressions or where the tokens go, keep the tracing. Add the watchdog because tracing is blind to the run that never started, and because an alert is not a brake on spend.

Can I not just build the absence check myself? You can, and I did, twice, badly. The hard parts are not the happy path: deduplicating alerts so one broken job does not send 40 messages, distinguishing a late run from a dead one, and making sure the checker cannot take down the job it watches. That is roughly the whole product.

Does it work for jobs that are not LLM calls at all? Yes. A backup script, a data sync or a deploy uses the same rv run wrapper or the same ping URL. The cost detectors simply stay quiet when there is no cost reported.

RunVouch

Related field notes


Try it: free for 3 agents · Docs: Claude Code · cron