Field notes · 2026-09-28 · RunVouch
Langfuse vs a watchdog: you don't need tracing to know your nightly agent is broken
Tracing tells you why a run went wrong. It cannot tell you the run never happened. Where Langfuse fits, where a watchdog fits, and how to run both.
I have run agents unattended for about two years: a crypto trading bot, a handful of nightly Claude Code jobs, a weekly build that regenerates a static site. In that time I have broken things in two clearly different ways. Sometimes a run did the wrong thing, and I needed to see exactly which tool call went sideways. Sometimes a run did nothing at all, and I found out four days later because a number on a page had stopped moving.
Tracing solves the first problem well. It does not solve the second, and that is not a flaw in tracing. It is a different question.
RunVouch is a dead man's switch, cost cap and outcome check for unattended AI agents: it alerts when a scheduled run does not happen, fails, produces no evidence, loops on the same tool call, or spends more than you allowed.
What Langfuse is good at
Langfuse is an open source LLM observability platform. You instrument your application and it records the run: the prompt that went out, the response that came back, token usage, latency, retrieval steps and tool executions. Instrumentation happens inside your code, through a drop-in client wrapper, a framework callback handler, decorators, or an OpenTelemetry span processor (Langfuse tracing docs). When an agent produced a wrong answer, or a chain got slow, or one step quietly burned most of the tokens, the trace is where you find out why.
I use that kind of data. It is the reason I would not tell anyone to drop tracing. But it requires the run to have happened, and to have gotten far enough to emit spans.
Is there a Langfuse alternative for cron jobs and nightly agents?
Honestly: the question is usually framed wrong. What people mean is "I want to know my 03:00 job ran and did not cost $40", and they reach for the tool they already know.
Langfuse can get part of the way there. Monitors combine a metric, a threshold and a window such as one hour, one day or one week, and route alerts to Slack, a webhook or a GitHub Action. There is even explicit no-data handling: you can treat null as zero, keep the previous severity, show a NO_DATA state, or notify after a sustained absence of data (Langfuse monitors and alerts). So an "our traces went quiet" alarm is buildable.
Here is what that still does not give you, and why I ended up building something separate:
- It is a query over a window, not a per-agent schedule contract. "This agent must check in every 24 hours" is a property of the agent, not a dashboard filter.
- A cron line that never fired, a laptop that was asleep, a container that was never scheduled, a wrapper script that died before the Python process started: no spans, and no way to distinguish "never started" from "started and hung".
- Nothing stops spend. An alert that a threshold was crossed arrives after the money is gone, and the next scheduled run starts anyway.
- A trace saying the model returned 9,000 tokens of confident prose is not evidence that the report file was written.
Tracing answers why, a watchdog answers whether
The split I use:
- Whether: did it run, did it finish, did it produce the artifact, did it stay inside budget. This has to work when the agent process is dead, so it cannot live inside the agent process.
- Why: which step, which prompt, which tool call, how many tokens. This has to live inside the run, because that is where the detail is.
Put the "whether" layer outside the thing it watches, and keep it boring. Mine is a small server with SQLite behind it, and the client fails open: if the monitoring endpoint is down, the wrapped job still runs. Monitoring that can break the job is worse than no monitoring.
Lightweight LLM monitoring without running a Clickhouse cluster
If you self-host Langfuse, the documented deployment needs Postgres for transactional data, Clickhouse for traces, observations and scores, Redis or Valkey for queues and cache, and S3 or another blob store for raw events and exports, alongside the web and worker containers (Langfuse self-hosting docs). For a team with real trace volume that is a reasonable bill. For one person with six cron jobs, it is a second system to keep alive, and the failure I actually care about is a job that did not run.
RunVouch is deliberately at the other end: one FastAPI server, one SQLite file, MIT licensed if you want to self-host it, EU hosted if you do not. Free for 3 agents, $9 for Solo, $29 for Team.
What Helicone's maintenance mode says about choosing a monitoring tool
Helicone moved into maintenance mode after being acquired by Mintlify, announced in March 2026, and Langfuse now publishes a migration guide that describes it plainly: services remain live for the foreseeable future, but the product is no longer actively developed, and the advice is to export your data early (Langfuse migration guide).
Two lessons I took from that. Keep the layer that must never lie small, portable and self-hostable, because migrating it should be an afternoon and not a quarter. And be suspicious of a setup where losing your observability vendor also means losing your only signal that a nightly job stopped working.
How I wire up both on one job
Declare the agent once, then wrap the command:
rv agent nightly-report --cadence 24h --cap-run-cost 2 --evidence rv run nightly-report --evidence-file out/report.html -- claude -p "$(cat brief.md)"
The cadence is the dead man's switch: no check-in inside 24 hours and I get a MISSED alert on Telegram, Slack or a webhook. The evidence file is the outcome check: exit code zero but an untouched output file is NO_EVIDENCE, not success. The cap is the brake: crossing it pauses the agent, and the next rv run is refused. It never kills a running process, because killing a half-finished trading or deploy step is its own kind of damage.
Inside Claude Code I add the plugin, which hangs off the SessionStart, PostToolUse and Stop hooks (Claude Code hooks reference) and reads tokens and cost from the transcript. There is an MCP server in the official registry as com.runvouch/runvouch, plus Python and Node clients. For anything that cannot host a client, an n8n branch, a Home Assistant automation, a shell one-liner on a NAS, each agent has a ping URL: /ping/<token> to report success, /ping/<token>/start, /ping/<token>/fail, or the exit code appended. Anything that can call a URL can check in.
The detectors that run on top of those check-ins: MISSED, FAILED, NO_EVIDENCE, RETRY_STORM (the same tool with identical input 8 times or more), BUDGET_RUN, BUDGET_DAY, DRIFT (runtime or cost far off the last 7 runs, by median absolute deviation) and STALLED. When one of those fires, that is when I open the trace.
FAQ
Should I replace Langfuse with a watchdog? No. If you are debugging agent quality, prompt regressions or where the tokens go, keep the tracing. Add the watchdog because tracing is blind to the run that never started, and because an alert is not a brake on spend.
Can I not just build the absence check myself? You can, and I did, twice, badly. The hard parts are not the happy path: deduplicating alerts so one broken job does not send 40 messages, distinguishing a late run from a dead one, and making sure the checker cannot take down the job it watches. That is roughly the whole product.
Does it work for jobs that are not LLM calls at all? Yes. A backup script, a data sync or a deploy uses the same rv run wrapper or the same ping URL. The cost detectors simply stay quiet when there is no cost reported.
RunVouch
Related field notes
- My Claude Code cron ran up $1,800 in two nights. The watchdog that stops it at $2
- Claude Code Routine failed silently? How to know your scheduled agent actually ran
- The dead man's switch for AI agents (and why a ping isn't enough)
Try it: free for 3 agents · Docs: Claude Code · cron