Pulse is a server-side Vintage Story mod that serves the server's own health numbers on a
Prometheus scrape endpoint. It runs on the dedicated server only, ships as a single dll with no
bundled dependencies, and does not talk to anything on its own: something has to come and read
/metrics. A separate optional mod pushes the same metrics over OTLP, described further down.
The documentation has its own site:
stratumserver.github.io/Pulse. New to Prometheus
and Grafana? Its Getting started page
walks through installing Pulse and getting your first dashboard, step by step. The same guide is
in this repository as docs/getting-started.md.
Grab both from the ModDB page or from GitHub releases.
- Install
- Configuration
- Scraping it
- Degraded mode
- Runtime metrics
- Attribution
- OTLP export
- Building and testing
- Where this is going
- License
First time with Prometheus and Grafana? docs/getting-started.md
covers install through your first dashboard, step by step.
Drop pulse_x.x.x.zip into your server's Mods/ folder and start the server. Add
pulseotlp_x.x.x.zip beside it if you want OTLP push as well; the base mod works on its own and
the OTLP one does not. On first boot Pulse writes ModConfig/pulse.json with its defaults:
{
"Enabled": true,
"Bind": "127.0.0.1",
"Port": 9464,
"RuntimeMetrics": true,
"ChunksRefreshSeconds": 30,
"Attribution": {
"Enabled": false,
"BurstTicks": 10,
"IntervalSeconds": 10
}
}Set Enabled to false and the mod loads but registers nothing at all: no tick listener, no
socket, no meter. RuntimeMetrics false drops the dotnet_* families and keeps the rest, which
is what you want if something else already collects them on that host. ChunksRefreshSeconds
is how often the loaded-chunk gauge is refreshed, and 30 is already fast for what that read
costs; lower it only if you know why. Attribution is the per-mod breakdown described further
down, off because it costs tick time.
Everything outside the Attribution block takes a server restart. The block itself does not:
/pulse reload applies it live, and /pulse attribution on and off switch it without touching
the file at all.
Upgrading does not mean editing the file by hand. Each mod checks its config file at startup and
writes back any key it knows about that the file is missing, with that key's default; the values
you already set are kept exactly as they are, and the log lists what was added. A key neither mod
recognises does not survive that rewrite, so it is reported as a warning instead of disappearing
quietly: usually it is a typo, and the setting you meant has been running on its default. A file
that already holds every key is not written at all, which matters if you mount ModConfig
read-only or keep it under version control.
Pulse and the OTLP mod each keep their settings in their own file under ModConfig/, written
with their defaults on first boot. Both files pick up new keys the same way an upgrade adds
them to a file that predates those keys; see the paragraph on that in Install.
A file that exists but will not parse is left exactly as it is: the mod logs the full path and
the parser's own message, and Pulse runs that session on its built-in defaults rather than
failing to start. Pulse OTLP does the same for pulse-otlp.json, except exporting stays off for
that session instead of falling back to its own default endpoint, and any value Newtonsoft
quoted in its message, a misconfigured Headers entry most often, is redacted before it reaches
the log.
ModConfig/pulse.json
| Key | Default | What it does | Live or restart |
|---|---|---|---|
Enabled |
true |
Turns the mod on. false loads it but registers nothing: no tick listener, no socket, no meter. |
Restart |
Bind |
"127.0.0.1" |
Address the metrics endpoint binds. See A word on the bind address. | Restart |
Port |
9464 |
Port the metrics endpoint listens on. If it is already taken, Pulse logs an error and runs without the endpoint. | Restart |
RuntimeMetrics |
true |
Serves the .NET runtime's own dotnet_* metrics alongside Pulse's. See Runtime metrics. |
Restart |
ChunksRefreshSeconds |
30 |
How often the loaded-chunk gauge, and the entity breakdown riding the same listener, are refreshed. Floored at 1 second, capped at one day (86400). | Restart |
Attribution.Enabled |
false |
Turns per-mod tick attribution on. See Attribution. | Live, via /pulse reload |
Attribution.BurstTicks |
10 |
Consecutive ticks profiled per burst. Clamped to 1 through 300. | Live, via /pulse reload |
Attribution.IntervalSeconds |
10 |
Seconds between the end of one burst and the start of the next. Floored at 1. | Live, via /pulse reload |
ModConfig/pulse-otlp.json
The OTLP mod has no reload command: every key below needs a restart to take effect.
| Key | Default | What it does | Live or restart |
|---|---|---|---|
Enabled |
true |
Turns OTLP export on. false keeps the mod loaded but exports nothing. |
Restart |
Endpoint |
"http://localhost:4318" |
Base address of the collector, without a signal path. Pulse appends /v1/metrics for http/protobuf; the exporter appends its own service path for grpc. |
Restart |
Protocol |
"http/protobuf" |
http/protobuf or grpc. Anything else logs a warning and falls back to http/protobuf. |
Restart |
Headers |
{} |
Headers sent with every export, for backend authentication. See OTLP export. | Restart |
IntervalSeconds |
60 |
Seconds between two exports. Floored at 5, capped at 86400 (24 hours). | Restart |
IncludeRuntimeMetrics |
true |
Adds the System.Runtime meter to what gets pushed. Independent of the base mod's RuntimeMetrics. |
Restart |
ServiceName |
"vintagestory" |
Sets the service.name resource attribute. A blank value falls back to vintagestory; OTEL_SERVICE_NAME, if set, overrides this key. |
Restart |
scrape_configs:
- job_name: vintagestory
static_configs:
- targets: ["127.0.0.1:9464"]GET /metrics returns the exposition text; every other path returns 404.
A ready-to-run Prometheus and Grafana pair lives in contrib/grafana; Prometheus alerting rules
calibrated to these thresholds live in contrib/alerts.
The exposition text is the contract: game panels can read /metrics directly instead of going
through Prometheus, which is how the first panel integration was built. Three things to know.
Each server instance runs its own Pulse on its own port, so a shared machine has one endpoint
per instance. The loopback bind covers a panel running on the same host; scraping from another
machine goes through a reverse proxy or a deliberate Bind change, as below. Any polling
cadence works, the endpoint is cheap to hit; existing metric families keep their names and
shapes, and anything breaking would be called out loudly in the changelog first.
The default binds loopback, which means only something running on the same host can scrape it.
That default is deliberate. A Vintage Story server is usually a public host, and the metrics
endpoint has no authentication of any kind, so widening Bind to 0.0.0.0 publishes your
player count and tick health to whoever asks. Anyone who can reach the port can also occupy its
(small, fixed) number of connection slots and blind your own scraper behind them, so a Bind
beyond loopback wants a firewall rule limiting the port to the scraper. Changing Bind is a
choice you should make on purpose, not a default you inherit.
If you need to scrape from elsewhere, the safest options leave Bind on loopback: tunnel to it,
or put a reverse proxy in front of it that only your scraper can reach. If you widen Bind
instead, add the firewall rule above.
Bind is a plain socket address, not a URL prefix. 0.0.0.0 binds every IPv4 interface, and
localhost binds the IPv4 loopback directly, so both localhost and 127.0.0.1 reach it,
whichever one a client's own name resolution tries first. All of this behaves the same on
Windows, Linux and macOS, and none of it needs administrator rights or a netsh URL reservation
on Windows. Binding 0.0.0.0 there can still prompt Windows Firewall to ask whether to allow
access, or be blocked outright by a service's default inbound rules; loopback never asks, since
nothing outside the machine is trying to reach it.
If the port is already taken, Pulse logs an error and carries on without the endpoint. The game server keeps running; you get no metrics until you fix the config.
The metric families it serves:
pulse_server_ticks_total(counter): server ticks processed since startup. Prometheusrate()over it is your TPS.pulse_server_tick_seconds(histogram): wall-clock seconds between consecutive ticks, measured with the mod's own stopwatch, in buckets from 25 ms to 1 s.pulse_players_online(gauge): players currently connected.pulse_entities_loaded(gauge): entities loaded in the world.pulse_server_tick_budget_seconds(gauge): the configured tick budget, sotick_seconds / tick_budget_secondsreads as saturation.pulse_worldgen_queue_columns(gauge): chunk columns waiting in the generation queue.pulse_worldgen_columns_generated_total(counter): chunk columns generated since startup.pulse_chunks_loaded(gauge): chunks loaded in the world, on a slow cadence of its own.pulse_log_entries_total{level}(counter): log entries by severity, one series each forwarning,errorandfatal.pulse_engine_warnings_total{kind}(counter): the engine's own health warnings, recognised by the text it logs.overloadis a tick past 500 ms,memoryis crossing 90% ofDieAboveMemoryUsageMb,suspend_timeoutis a server suspend that gave up waiting for a thread, andautosave_iois an autosave arriving while the previous one is still writing.pulse_server_uptime_seconds(gauge): seconds the server has been ticking. This is the engine's own unpaused clock, so it stops during a save and is not process uptime.pulse_entities_by_code{code}(gauge): loaded entities by entity code, the ten most numerous plus anotherbucket for the rest. Refreshed on the same slow cadence as the chunk count.pulse_player_ping_seconds{stat}(gauge): round trip time to the players online, asavgandmax. Both read 0 with nobody connected.pulse_player_deaths_total(counter): player deaths since startup.pulse_server_suspends_totalandpulse_server_suspend_seconds_total(counters): how often the server suspended ticking and how long it spent suspended. Every autosave is one of these, and the seconds are the pause players actually feel.pulse_network_sent_bytes_totalandpulse_network_received_bytes_total(counters): bytes over the main TCP channel. UDP is not in these two; the public server API does not report it.
Six more come from the engine's own accounting, which no public API exposes. See the note on degraded mode below for what happens when they are unavailable.
pulse_server_tick_busy_seconds(gauge): the average time one tick spent working, sleep excluded, over the engine's last completed two-second window. This is the number/statsprints, and the only view of headroom below the tick budget that exists at all. The engine measures in whole milliseconds, so an idle server legitimately averages zero.pulse_network_packets_per_second{channel}andpulse_network_bytes_per_second{channel}(gauges): traffic rates over that same window, splittcpandudp. Gauges rather than counters because the engine zeroes its window rather than accumulating it; a window cut short by a suspend reads low for one sample.pulse_connection_queue_clients(gauge): clients waiting because the server is full.pulse_network_udp_sent_bytes_totalandpulse_network_udp_received_bytes_total(counters): the UDP totals missing from the two public byte counters above.
Four more answer "which mod is eating the tick", and only when you turn them on. They have a section of their own further down.
The tick period is measured rather than taken from the value the engine hands tick listeners, because that one is rounded to whole milliseconds. Overruns still land exactly: once a tick's work exceeds the budget the engine's throttle sleep is zero, and the period is the busy time. Most gauges come from a snapshot the tick listener refreshes about once a second on the server's main thread, so a scrape never touches live world state.
Loaded chunks are the exception. The engine offers no cheap count, and the one accessor that
exists clones the entire loaded-chunk dictionary under the chunk lock, so that gauge gets its own
listener at ChunksRefreshSeconds and reads 0 until the first refresh. The entity breakdown
rides that same slow listener, but it is not the same kind of gauge: pulse_entities_by_code is
recorded only from the listener callback, with no observable default behind it, so it is absent
from the exposition entirely until that first refresh, not present at 0. The event-driven
counters do not ride the tick listener either: columns generated is incremented from
MapChunkGeneration, which fires on the worldgen thread, and the log counters from
Logger.EntryAdded, which fires on whichever thread wrote the line. Both handlers classify and
increment, and nothing else.
Six of those families come from a place the modding API does not reach. Tick busy time, the
per-window packet and byte counts, the connection queue depth and the UDP byte totals are all
measured by the engine on concrete types inside VintagestoryLib.dll, which the game's authors
change freely between versions and make no promises about. Pulse casts the server world to
Vintagestory.Server.ServerMain once, at startup, in a try/catch, and every read of those types
lives in one small class.
If that cast ever stops working, Pulse logs one warning naming what went wrong and carries on.
The six families are simply absent from /metrics rather than present and lying, and every other
metric on this page keeps being served exactly as before, including the tick period histogram,
which still catches every overrun: once a tick's work exceeds the budget the engine's throttle
sleep is zero and the period is the busy time. What you lose is the view of headroom on a healthy
server. One warning, no retry loop, no log spam, and nothing that can stop a game server.
The compile-time reference to VintagestoryLib.dll is deliberate for the same reason. A game
version that renames or moves ServerMain breaks the Pulse build, loudly, before a release goes
out, instead of shipping a mod that quietly serves six families fewer.
With RuntimeMetrics left on, the .NET runtime's own System.Runtime meter is served
alongside Pulse's, as dotnet_* families: GC collections and pause time, heap size and
fragmentation by generation, working set, CPU time by mode, JIT, thread pool, lock contention,
loaded assemblies. None of it is instrumented here. The runtime publishes the meter, and the
writer renames each instrument the way Prometheus's otlptranslator does, the library Prometheus's
own OTLP receiver, Mimir and Grafana Cloud use to turn an OTLP instrument into a Prometheus name:
dots become underscores, the instrument's unit becomes a trailing word unless the name already
contains it as a word, and a monotonic counter's name ends in _total, moved there rather than
duplicated if the name already spells "total" somewhere. dotnet.gc.collections is served as
dotnet_gc_collections_total; dotnet.process.memory.working_set, a gauge in bytes, is served as
dotnet_process_memory_working_set_bytes; dotnet.gc.heap.total_allocated, a counter also in
bytes, is served as dotnet_gc_heap_allocated_bytes_total rather than the doubled
..._total_allocated_bytes_total. This is the same name Grafana derives when it translates the
OTLP export, so a dashboard or alert built against a server scraped over OTLP through Grafana
Cloud reads Pulse's own /metrics without translation too.
Nine of these families moved to this spelling in 0.2, to line up with that translation; see the
changelog for the full old to new list if you have a dashboard or alert built against the earlier
names. pulse_* families are unaffected.
Pulse renders the shape each instrument declares, including where that is arguable.
dotnet_thread_pool_thread_count_total is typed as a counter because the runtime publishes it
as an ObservableCounter, even though the number goes down as often as up. Second-guessing the
framework here would only make the series harder to correlate with any other .NET exporter.
Tick busy time tells you the server is working hard. Attribution tells you what it is working on. Turned on, it reports a per-mod share of the main thread, on a continuous series you can graph and alert on, rather than in a one-off profiling report.
It is off by default, because it is not free. Add an Attribution block to ModConfig/pulse.json:
{
"Attribution": {
"Enabled": true,
"BurstTicks": 10,
"IntervalSeconds": 10
}
}BurstTicks is how many consecutive ticks each measurement covers, IntervalSeconds how long the
server runs unmeasured between two of them. The defaults measure about one tick in thirty. Both are
clamped on read: at least a second between bursts, at most 300 ticks in one.
The moment you want attribution is usually while the server is struggling, and restarting it throws
away the thing you wanted to look at. So the config file is not the only way in. Four commands, all
behind the controlserver privilege, so admins and a panel console can run them:
/pulse attribution onstarts the duty cycle straight away, on whateverBurstTicksandIntervalSecondsare in force. The first burst lands one interval later./pulse attribution offstops it and puts the engine's frame profiler back down, unless the engine's own/debug logticksstill wants it running. A burst in progress is dropped rather than published half-measured./pulse attribution statusreports whether it is running, the cycle it is using, how many ticks it has profiled and whether it is inside a burst right now./pulse reloadre-readspulse.jsonand applies theAttributionblock live. The reply names any other key whose value in the file has drifted from what the server is running, since those still need a restart, and a file that does not parse changes nothing at all.
on and off act on the running server and never write pulse.json, which is deliberate: a ten
minute look should not become permanent because somebody forgot to turn it off. Restart the server
and the file decides again. To make a change stick, edit the file and either restart or run
/pulse reload.
This works even on a server that booted with Attribution.Enabled false. Pulse registers the four
families and primes the engine's profiler at startup either way, because a profiler switched on
part-way through a tick that has never completed one takes the server down with it (there is more
on that below). Priming costs two profiled ticks at boot and nothing after; an instrument nothing
has recorded into is not a series, so an idle server serves exactly the exposition it did before.
Four families appear once it is on:
pulse_mod_tick_share{modid}(gauge): the fraction of profiled main-thread busy time that went to one mod over the last completed burst. The shares add up to 1 across everymodid, including the two Pulse adds:enginefor the server's own systems and for the time no marker named, andunattributedfor work that was marked but that no loaded mod claims.pulse_mod_tick_seconds_total{modid}(counter): main-thread seconds attributed to one mod. Sampled, not total: this is time measured inside the bursts, not time since startup. Divide by the tick counter below to compare two servers, or takerate()of it againstrate(pulse_attribution_ticks_total)for seconds per profiled tick.pulse_attribution_ticks_total(counter): ticks actually profiled, which is what makes the sampled seconds mean anything.pulse_attribution_dropped_samples_total(counter): profiler readings thrown away because they overflowed. The engine accumulates each marker's time into a 32 bit counter of stopwatch ticks, which wraps negative somewhere past two seconds inside a single tick. A wrapped reading is not a large number, it is garbage, so it is dropped and counted here instead of being published as data. Anything but a flat zero means the server had a tick so bad that a single marker ran for over two seconds.
The engine already contains a per-mod tick attributor and simply never switches it on. With its frame profiler enabled, the server stamps a marker after every game tick listener, every delayed callback and every main-thread entity behaviour, keyed by the type that declared the handler or by the behaviour's registered code. Pulse turns the profiler on for a burst, reads the tree the tick left behind, maps each key back to a mod through the mod loader, and turns it off again. No Harmony, no engine patch, no bundled dependency.
The cost is measured, not estimated from a mark count. Pulse.Scenarios/AttributionCostScenarios.cs
joins a test player, spawns four thousand chickens (dense cluster and, separately, spread across
the loaded area, to tell density from entity count apart), and reads tick busy time with
World.MeasureTicks across off, on, off, on, off, on, off: four off windows bracketing three on
windows, each on compared against the mean of the two off windows next to it rather than just the
one before it, so a baseline that drifts across the run cannot bias every delta the same direction.
On a quiet run, both load shapes agreed: the off baseline wobbled by a couple of milliseconds with
no clear drift, and a profiled tick's own share came out to about 8.75 ms, roughly 26% of the 33 ms
budget. A second scenario splits that further, forcing the engine's frame profiler on without
letting Pulse fold what it records: about 9 ms (27% of budget) is the engine's own cost of writing
the marks. Pulse's own share, read from its own attribution at the stopwatch resolution that
measures at rather than from a millisecond-rounded difference too small for that resolution to see,
comes to about 0.02 ms, under a tenth of a percent of the budget: there is nothing here for Pulse
itself to usefully optimise. At the shipped default (10 ticks every 10 seconds) the measured share
blends down to about 0.9% of the budget, a slow window of roughly a third of a second every 10
seconds. The previous default was 30 ticks; the same measurement puts that at about 2.5%
amortised, and the shorter burst is what shipped once the real cost of the longer one was known.
All of these replace the mark-count estimate this section used to carry (2.8% burst, 0.3%
amortised at the old default), measured too low, likely because a dictionary write and a clock
read cost more in practice than its per-operation guess. Markers scale with loaded entities times
their behaviours, not with how many mods you run, so raising BurstTicks or lowering
IntervalSeconds moves the amortised share in the obvious direction, and on an idle server it is
nothing at all. Run it yourself, with VINTAGE_STORY set: PULSE_MEASURE_ATTRIBUTION_COST=1 dotnet test Pulse.Scenarios --filter "Category=Cost"; CI filters the Cost trait out, and locally
it is a no-op unless that variable is set, because spawning four thousand entities, three times
over, is slow.
One visible side effect: the engine logs "Over 400ms tick. Skipping N physics ticks" only while its
frame profiler is on, and Pulse is what turns it on. It does so for the first tick after every
start, attribution enabled or not, to prime the profiler so attribution can be switched on later
without a restart. That first tick loads the spawn area and usually runs long, so one such line
at startup is normal and harmless; it also counts once in pulse_log_entries_total{level="warning"}.
During attribution bursts the same line appears whenever physics falls behind: that is the engine
reporting a real condition it otherwise keeps to itself.
Say this out loud before reading a dashboard built on it.
Broadcast events carry no markers. Roughly forty of them, PlayerJoin, DidBreakBlock,
OnEntityDeath and the rest, are plain C# events the engine invokes without timing. A mod that
does all its work in an event handler shows up as a rounding error here, and the time it spends
lands in the engine bucket. The listener-and-behaviour half is what this measures.
It is a main-thread share, not a total. Entity behaviours that declare themselves thread-safe run across several threads, and only the main thread's slice is marked. A mod whose behaviour is thread-safe therefore reads low, by roughly the thread count.
Mapping is by assembly. A mod that ships several dlls only has the one its ModSystem lives in
claimed, so a listener registered from a side library reads as unattributed. So does a handler
on a static method, which the engine marks with no identity at all.
And it is a sample. Ten ticks every ten seconds describe a steady server well and a spiky one badly. The share is an average over the burst, so a mod that stalls for 200 ms once a minute may well be profiled during a quiet stretch and read as harmless.
If the numbers matter enough to act on, this is a first pass that says which mod to look at, not a call tree. Lithos Probe's sampling profiler is the tool for the second pass.
OTLP is the push counterpart to the scrape endpoint: instead of waiting for Prometheus to come
and read /metrics, the server sends its metrics to a collector on a timer, in the wire format
every major observability backend accepts. Grafana Cloud, Honeycomb, Datadog, New Relic and an
otel-collector you run yourself all take the same payload.
It ships as a second mod, pulseotlp_x.x.x.zip, and both zips go in Mods/. The base mod stays
a single dll with no dependencies; the OTLP one carries the OpenTelemetry SDK and its
Microsoft.Extensions.* fan-out, eighteen dlls in all. That split is not tidiness. The game's
mod loader puts every root-level dll of every mod into one shared assembly context with no
version arbitration, so a bundled dependency is a collision risk against every other mod on the
server, and a server that does not push metrics should not be paying it.
The two mods share no code. Pulse.Otlp.dll has no reference to Pulse.dll; it subscribes to
the meter named Pulse.Server, which is all System.Diagnostics.Metrics needs, and
modinfo.json declares the dependency so the loader guarantees the base mod is there first.
On first boot it writes ModConfig/pulse-otlp.json:
{
"Enabled": true,
"Endpoint": "http://localhost:4318",
"Protocol": "http/protobuf",
"Headers": {},
"IntervalSeconds": 60,
"IncludeRuntimeMetrics": true,
"ServiceName": "vintagestory"
}Those defaults suit a collector running on the same host. Endpoint is the base address, without
a signal path: Pulse appends /v1/metrics for http/protobuf and leaves it alone for grpc,
where the exporter appends its own service path. Protocol takes the two names the OTLP
specification defines, http/protobuf and grpc; anything else logs a warning and falls back to
http/protobuf rather than leaving you with no export at all. IncludeRuntimeMetrics adds the
System.Runtime meter to what gets pushed, and it is separate from the base mod's
RuntimeMetrics flag, so you can serve the dotnet_* families locally and not ship them, or the
other way round.
ServiceName sets the service.name resource attribute, which is how a backend receiving
metrics from more than one server tells them apart: grouping, filtering and dashboard variables
are usually keyed off it. The OTEL_SERVICE_NAME environment variable, the ecosystem's standard
override, takes precedence over this key when it is set. For a stable instance label, pair
OTEL_SERVICE_NAME (not this key) with OTEL_RESOURCE_ATTRIBUTES=service.instance.id=<id>:
without OTEL_SERVICE_NAME, a freshly generated id silently overrides that variable on every
restart.
IntervalSeconds is floored at 5 and capped at 86400 (24 hours), the cap there so a config typo
several digits too long cannot overflow the millisecond count it is converted to. Sixty is the
OTLP default and the right answer for almost everyone: the interval also decides how often every
observable gauge is polled, and the loaded-chunk read behind one of them is not free.
For a local collector, the whole config is the endpoint:
{
"Endpoint": "http://localhost:4318",
"Protocol": "http/protobuf"
}A hosted backend wants an auth header. Grafana Cloud's OTLP endpoint takes HTTP basic auth, with the instance ID as the user and an access policy token as the password:
{
"Endpoint": "https://otlp-gateway-prod-eu-west-2.grafana.net/otlp",
"Protocol": "http/protobuf",
"Headers": {
"Authorization": "Basic MTIzNDU2OmdsY19leGFtcGxldG9rZW4="
},
"IntervalSeconds": 60
}Headers goes out with every export, so anything a backend accepts works the same way:
x-honeycomb-team for Honeycomb, api-key for New Relic, x-scope-orgid for a multi-tenant
Mimir. Values are percent-encoded on the way into the exporter, which the OTLP header format
expects, so a base64 token with +, / and = in it needs no special handling. A literal comma
in a header value does not survive the trip: the exporter unescapes the whole header string before
splitting it on commas, so encoding it going in does not stop it from being read as a separator
coming out. Pulse checks every header for this, and for two names that collide once leading and
trailing whitespace is trimmed off, before ever handing them to the exporter, and refuses to start
exporting rather than let either reach it: the server log names the offending header, never its
value.
pulse-otlp.json holds a credential. It sits in ModConfig/ in plain text, with whatever
permissions your server's umask gave it. On a shared or rented host, chmod 600 it and make sure
it is owned by the account the server runs as. It is also worth keeping out of any config backup
you push somewhere public.
A collector that is down, refusing, or answering 401 still costs you nothing on the game side: the
OpenTelemetry SDK exports from its own background thread and the tick loop never sees the
failure. It no longer stays invisible, though. Pulse OTLP listens to the SDK's own diagnostic
event source and turns the first failure of each kind into one line in the server log, repeated at
most every ten minutes and logged at Warning rather than Error so a struggling backend can never
count toward DieAboveErrorCount:
Pulse OTLP export to https://otlp-gateway-prod-eu-west-2.grafana.net/otlp/v1/metrics failed:
Response status code does not indicate success: 401 (Unauthorized). The backend answered:
{"status":"error","error":"authentication error: invalid token"} Metrics are not reaching the
backend; check Endpoint and Headers in pulse-otlp.json. This is logged again at most every 10
minutes.
A matching line reports the first successful export after a failure, so recovery shows up too, and a healthy server that has never failed still logs exactly one such line, at Notification rather than Warning, right after its first delivery.
Neither line is meant to carry a header value, or the query string or userinfo half of Endpoint
either, for a backend that authenticates a signed URL that way instead of through a header. Before
a line is queued, each of those values, at least 6 characters long (shorter than that reads as an
ordinary id, not a credential), the credential half of it when the value has a "scheme credential"
shape (a Bearer token echoed without its "Bearer ", say), and the JSON-escaped form of both, are
matched case-insensitively and redacted out of the backend's answer and out of a gRPC failure's
status detail, longest value first so a short one can never land inside a longer one's own match.
Anything else shaped like a bearer or basic credential of at least 8 characters is redacted too,
whether or not it matches a configured value. This is not exhaustive: a backend that transforms a
secret some other way, hashing it or splitting it across two fields, could still get it into the
log, so treat the log itself as sensitive before sharing it regardless. The backend's answer, once
redacted, is clipped to 200 characters.
A malformed Endpoint, and a Headers entry with a comma in its value or a name that collides
with another once trimmed, are cases Pulse checks itself before the exporter is ever built: it
logs one error, naming the problem and never the value, and registers nothing. An IntervalSeconds
so large it would once have overflowed the millisecond conversion is not one of those cases any
more: it is silently clamped to 86,400 seconds (24 hours) and exported at that rate instead, with
no error at all. A Headers shape neither check above names, such as a comma inside a header
name rather than its value, or a header called User-Agent (which collides with one the
exporter sets on its own), still reaches the OpenTelemetry SDK's own option validation, which
throws; Pulse catches that too, so nothing crashes and nothing is exported, but the log line only
names the exception type, not the header. For anything these lines do not explain, the SDK's own,
far more verbose self-diagnostics turn on by dropping an OTEL_DIAGNOSTICS.json file next to the
server.
You need the .NET 10 SDK and a Vintage Story 1.22.x install, with VINTAGE_STORY pointing at
the folder that holds VintagestoryAPI.dll and VintagestoryLib.dll (the .pdb next to the
first is required too, or the engine's logger crashes at boot). Both dlls are compile-time
references only; neither is copied into the mod, which still ships as one file.
export VINTAGE_STORY=/path/to/vintagestory
dotnet build Pulse.slnx -c Release
dotnet test # unit tests, then the Atlas scenarios
dotnet build Pulse/Pulse.csproj -c Release -t:PackageMod # artifacts/pulse_x.x.x.zip
dotnet build Pulse.Otlp/Pulse.Otlp.csproj -c Release -t:PackageMod # artifacts/pulseotlp_x.x.x.zipThe scenarios in Pulse.Scenarios boot a real headless server in-process through
Atlas, load the mod, and scrape it over HTTP for real. The
atlas CLI runs the same assembly without VSTest, which is faster to iterate against:
atlas run Pulse.Scenarios/bin/Release/net10.0/Pulse.Scenarios.dll
atlas run Pulse.Otlp.Scenarios/bin/Release/net10.0/Pulse.Otlp.Scenarios.dllPulse.Otlp.Scenarios is a separate project because it stages both mods, laid out exactly as
their zips are, and the base suite's staging should stay as it is. It stands up a fake collector,
points the mod at it, runs the world, and asserts on the protobuf that arrives. There is one
collector per protocol the config accepts: an HttpListener for http/protobuf, and for grpc a
small HTTP/2 server, since a gRPC client wants its own service path, a length-prefixed message
and a status trailer before it calls an export delivered.
Unit tests in Pulse.Tests cover the aggregator, the exposition writer, the log classifier and
the small classes behind the wave of engine and world metrics: the busy-time average, the ping
aggregates, the entity top-ten with its series retirement rule, and the suspend window. None of
them needs a server. Pulse.Otlp.Tests covers the config translation, which is where the OTLP
mod's only non-obvious logic lives. CI also runs both unit suites on Windows. The scenarios stay
on Linux, since they boot a server build made for it.
Mutation testing runs at two depths. tools/mutation-check.sh applies ninety-one representative
mutations one at a time and requires the suite to fail on every one; CI runs it on every code
change, deterministic and done in about six minutes. .github/workflows/mutation.yml runs
dotnet-stryker incrementally on pull requests into dev touching Pulse/: it mutates only the
files the pull request changed and reports without gating. The full run mutates the whole project
except the files that only run under a live server, scored 94 percent at the 0.2.0 release, fails
below 83, and runs weekly and on workflow_dispatch. Pulse.Otlp's own Stryker lane is
workflow_dispatch only, because Stryker launches the wrong project's test host for it and the
score swings too much between runs to gate a pull request on; it scored 81 percent at 0.2.0 and
fails below 70.
Static analysis runs on SonarCloud
for pushes to dev and pull requests into it. The job waits for the quality gate, so a failing
gate turns the check red instead of sitting unnoticed on SonarCloud's side. The gate judges new
code only: an A rating for reliability, security and maintainability, at least 80 percent
coverage, at most 3 percent duplication, and every security hotspot reviewed. The coverage badge
at the top of this page comes from the same job, which runs the two unit suites and the base
scenarios under coverlet. Five files are left out of that figure: the two ModSystems, the engine
probe, the attribution probe and the live-server half of the attribution metrics. The game's
loader loads the staged dll outside coverlet's instrumentation, so nothing a scenario executes in
them can reach a coverage report. The scenarios are still what tests them.
The documentation site is built by docs/site/build.mjs from the README, the getting-started
guide, the two contrib READMEs, the alert rules, the dashboard JSON and the changelog. It also
reads both modinfo.json, NOTICE and the replies in Pulse/PulseCommands.cs, and it stops when
any of these changes in a way a page depends on: a renamed section, a reworded /pulse reply. So
a pull request that touches no document can still fail the pages workflow, which runs the build
on every pull request and deploys the result from main. The build needs Node 22 or newer and
has two dependencies.
cd docs/site
npm ci
npm test # the build rules and the logic of the client scripts
npm run build # writes docs/site/dist
npm run serve # serves it on http://127.0.0.1:4173 (needs Python 3)The survey in docs/metrics-feasibility.md lists what is measurable on this engine and what is
not. Save duration is the honest gap: the world-save event fires before any writing happens and
there is no completion signal, so Pulse reports the suspend window instead of inventing a
number. Per-player traffic does not exist anywhere in the engine, not even internally.
Traces and logs over OTLP would reuse most of the exporter mod, but neither has an obvious consumer on a game server yet, so they stay unbuilt until someone asks.