|
|
2 kuukautta sitten | |
|---|---|---|
| .. | ||
| alertmanager | 3 kuukautta sitten | |
| grafana | 3 kuukautta sitten | |
| loki | 3 kuukautta sitten | |
| prometheus | 2 kuukautta sitten | |
| promtail | 3 kuukautta sitten | |
| README.md | 3 kuukautta sitten | |
| compose.yml | 2 kuukautta sitten | |
| postgres-exporter.compose.yml | 3 kuukautta sitten | |
phase 0 close-out — grafana + loki + prometheus + promtail on a single dokploy compose project. lives on the same docker network as apps/api, apps/runtime, apps/realtime so prometheus scrapes container names directly. grafana is exposed via traefik behind cloudflare access; everything else is internal-only.
infra/observability/
├── compose.yml # the stack
├── loki/config.yaml
├── promtail/config.yaml
├── prometheus/
│ ├── prometheus.yml
│ └── rules/services.yml
├── grafana/
│ ├── grafana.ini
│ ├── provisioning/datasources/datasources.yaml
│ ├── provisioning/dashboards/dashboards.yaml
│ └── dashboards/
│ ├── api-requests.json
│ ├── runtime-invocations.json
│ ├── realtime-subs.json
│ └── postgres-health.json
└── README.md (this file)
prometheus scrape targets in prometheus.yml:
| service | /metrics | metrics emitted today | needed for full dashboards |
|---|---|---|---|
api |
✅ shipped | http_requests_total{method,status,route}, http_request_duration_ms_bucket{method,route,le} |
already covered — see apps/api/src/lib/metrics.ts + apps/api/src/middleware/metrics.ts |
runtime |
✅ shipped | pool gauges, cold-start histogram, isolate kill counter | already covered — see apps/runtime/src/metrics.ts |
realtime |
✅ shipped | briven_realtime_subscriptions_active, briven_realtime_channels_active, briven_realtime_notifies_total, briven_realtime_reinvoke_total{outcome} |
already covered — see apps/realtime/src/metrics.ts |
postgres |
⚙️ template ready | none until deployed | merge ./postgres-exporter.compose.yml into the data-plane compose project + create a briven_metrics role (GRANT pg_monitor) + set POSTGRES_EXPORTER_DATA_SOURCE_NAME. exporter binds to postgres-exporter:9187, already in prometheus.yml. |
dashboards reference these metric names and will show "no data" until the corresponding service starts emitting them. that's deliberate — the dashboards land first so adding a metric is a one-line change.
logs are scraped by promtail using file-based discovery against /var/lib/docker/containers/*/*-json.log (read-only bind-mount). per DOCKER.md we never touch /var/run/docker.sock or docker_sd_configs — the daemon is shared with other projects on the host and every API call goes through internal locks. promtail reads the json log files directly, unwraps the docker envelope, then parses the service field from each briven service's own JSON log line (emitted by @briven/shared/observability). non-briven containers ship through with service absent and Loki queries scope by {service=~"api|runtime|realtime|web|docs"}.
# on the kvm running the briven docker network:
cd infra/observability
docker network create briven 2>/dev/null || true # idempotent
docker compose up -d
verify:
curl -fsS http://localhost:9090/-/ready # prometheus ready
curl -fsS http://localhost:3100/ready # loki ready
docker compose ps # all four containers up
then point cloudflare access at grafana.briven.tech (the traefik labels in compose.yml cover the rest).
phase 0 ships single-binary loki + local-disk storage. when log ingest exceeds ~50gb/day or query latency degrades:
common.storage.filesystem for s3 (minio or cloudflare r2)the dashboards and scrape configs don't change.
prometheus/rules/services.yml defines ServiceDown, HighErrorRate, PostgresConnectionsHigh. they route through alertmanager → benjojo/alertmanager-discord bridge → discord:
severity=critical|warning → #briven-alerts (page-worthy)#briven-deploysoperator env on the dokploy project:
DISCORD_WEBHOOK_ALERTS=https://discord.com/api/webhooks/…
DISCORD_WEBHOOK_DEPLOYS=https://discord.com/api/webhooks/…
create the channels in your discord server, generate webhook URLs, paste into the dokploy env. each bridge instance holds one URL; rotation = update env + restart that one service. existing alert state lives in the alertmanager_data volume — keep it across redeploys so silences carry through.
smoke test (drop a rule's for: to 0s temporarily, observe a message land, revert):
docker compose exec prometheus promtool query instant 'up'
docker compose restart prometheus
# … wait for the rule to fire …
# then revert the rule edit and `docker compose restart prometheus` again.
uids stay stable across redeploys — bookmark them.
briven-api-requests — request rate, p99 latency, 5xx error rate, recent error logsbriven-runtime-invocations — invocations/sec by project, cold-start p50/p99, active isolates, kills by reasonbriven-realtime-subs — active subscriptions, LISTEN channels, NOTIFY rate, listener reconnect eventsbriven-postgres-health — connections, transactions/sec, deadlocks, longest query, db size by schema