Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
39 commits
Select commit Hold shift + click to select a range
3022b92
feat(config): choose each layer's implementation at boot
EricAndrechek Sep 25, 2026
fde17ba
docs(config): say coord.backend is reserved; sync the boot-config lists
EricAndrechek Sep 25, 2026
f129d57
docs(config): no backend has a sub-block yet; index backends.go in AG…
EricAndrechek Sep 25, 2026
67df317
perf(cache): flat version index, pruned per tenant
EricAndrechek Sep 25, 2026
ed8022b
docs(cache): name the cache prune hook; pin LocalCache as a pruner
EricAndrechek Sep 25, 2026
44cec4d
Merge remote-tracking branch 'origin/feat/boot-backends' into feat/ca…
EricAndrechek Sep 25, 2026
85803f0
Merge remote-tracking branch 'origin/feat/cache-flat-versions' into f…
EricAndrechek Sep 25, 2026
c068084
feat(config): cache.backend=redis selects the shared cache
EricAndrechek Sep 25, 2026
13c6cb3
fix(cache): review round — e2e gate, trust boundary, #386, addr spaces
EricAndrechek Sep 25, 2026
ca16db5
Merge origin/feat/cache-snapshot into feat/cache-flat-versions
EricAndrechek Sep 25, 2026
ee51e84
test(app): assert the wired cache is a pruner before swapping it
EricAndrechek Sep 26, 2026
fcf1e89
Merge main (5004cd23) into feat/cache-redis-wiring
EricAndrechek Sep 26, 2026
9a4888a
fix(config): cache.redis defaults in defaults(); 0 never compresses
EricAndrechek Sep 26, 2026
6fa1ae6
fix(cache): give the two-instance test its roles; redis is the shared…
EricAndrechek Sep 26, 2026
712b1db
fix(config): refuse cache.redis.mode=sentinel until #656
EricAndrechek Sep 26, 2026
a7ef71f
fix(config): refuse a URL-style cache.redis address without echoing it
EricAndrechek Sep 26, 2026
de27b60
fix(config): cap cache.redis.dial_timeout at 2s
EricAndrechek Sep 26, 2026
8189783
test(cache): the local backend's insert-invalidates lifecycle, end to…
EricAndrechek Sep 26, 2026
5e84762
fix(ingest): log an invalidation that did not land at WARN
EricAndrechek Sep 26, 2026
ef59d85
Merge remote-tracking branch 'origin/feat/cache-snapshot' into feat/c…
EricAndrechek Sep 26, 2026
4df73f9
docs(cache): clarify which bumps a no-index tenant absorbs
EricAndrechek Sep 26, 2026
95575f2
docs(cache): CHANGELOG says #262 is part-fixed, not fixed
EricAndrechek Sep 26, 2026
e0d7018
Merge feat/cache-flat-versions (#621) into feat/cache-redis (#626)
EricAndrechek Sep 26, 2026
d60615b
Merge the new #626 and #621 heads into feat/cache-redis-wiring
EricAndrechek Sep 26, 2026
5b31ce6
fix(config): cap cache.redis timeout and dial_timeout at 1s
EricAndrechek Sep 26, 2026
41ebc66
fix(cache): review round 1 — one standalone address, e2e stack docs, …
EricAndrechek Sep 26, 2026
64f09e8
docs(dev): list Redis among the cache backend's integration servers
EricAndrechek Sep 26, 2026
558bb23
fix(cache): pin the no-index bump test and fix a "query key" mixup
EricAndrechek Sep 26, 2026
5a72724
fix(cache): review round 2 — stale docs claims, redisConfig field pin
EricAndrechek Sep 26, 2026
568f2dd
docs(cache): version_manager.go's own doc carries the query-key/entry…
EricAndrechek Sep 26, 2026
277b5b1
Merge remote-tracking branch 'origin/feat/cache-snapshot' into feat/c…
EricAndrechek Sep 26, 2026
c42b6db
docs(cache): fix the one query-key/entry-key spot the terminology pas…
EricAndrechek Sep 26, 2026
11a1fc6
Merge feat/cache-flat-versions (#621) into feat/cache-redis (#626)
EricAndrechek Sep 26, 2026
f1ab891
Merge the new #626 and #621 heads into feat/cache-redis-wiring
EricAndrechek Sep 26, 2026
ce0da48
docs(cache): match configuration.mdx and deployment.md to #626's fixes
EricAndrechek Sep 26, 2026
1058677
Merge the new #626 head (0eea1cbc) into feat/cache-redis-wiring
EricAndrechek Sep 26, 2026
ecd21c0
docs(cache): match deployment.md to #626's probe-fix (0eea1cbc)
EricAndrechek Sep 26, 2026
0281ca3
docs(cache): correct the Close bound, deferred count and probe wording
EricAndrechek Sep 26, 2026
28d4537
Merge the new #626 head (aad9b1e5) into feat/cache-redis-wiring
EricAndrechek Sep 26, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,7 @@ Break one of these knowingly or not at all.

2. **A dedicated `coverage` job applies the consolidated gate, polling — not `needs`-ing — the suites.** Each suite (`unit`, `integration`, `e2e`) runs with `COV_DEFER=1` and uploads a `coverage-<suite>` fragment; the `coverage` job runs `make cov` (merge + every threshold gate) over all three — exactly like local `make ci`'s final step. Keeping it a separate job (not folded into e2e's tail) decouples the gate result from the e2e suite's pass/fail. Crucially it is `needs: changes` **only, not the suites**: a `needs` edge is a *scheduling* barrier — GitHub won't pick up a runner, check out, restore caches, or `pnpm install` until the needed jobs finish — so needing the suites would serialize this job's ~50s of setup onto the critical path after the last suite, for nothing (the setup doesn't depend on their results). Instead it starts at run creation, runs its setup in parallel with the suites, and blocks only at the merge by polling for the three fragments with [`scripts/ci/wait-artifact.sh`](../../scripts/ci/wait-artifact.sh) (fails fast if a producer concluded without producing). Tail on the critical path: ~10s, not ~50s. **The aggregator and `docs-deploy` must keep `coverage` *and* every suite in their `needs`** — the suites directly (a suite failure must red the gate even though `coverage` no longer needs them), and `coverage` (else a coverage-gate failure wouldn't block merge or a prod deploy).

3. **e2e builds its own inputs and mirrors local `make test-e2e`.** It compiles the SDK dist + cover binary itself (`make -j test-e2e`, warm per-suffix cache) rather than waiting on a builder job, and runs the suite exactly as a developer does — one orchestrator, one ClickHouse testcontainer, sequential files. The ClickHouse image pulls in the background while caches restore (also in the integration job).
3. **e2e builds its own inputs and mirrors local `make test-e2e`.** It compiles the SDK dist + cover binary itself (`make -j test-e2e`, warm per-suffix cache) rather than waiting on a builder job, and runs the suite exactly as a developer does — one orchestrator, one ClickHouse and one Redis testcontainer, sequential files. The ClickHouse image pulls in the background while caches restore (also in the integration job).

4. **One change classifier, split into a pure core + a CI wrapper.** The pure allowlist — file list on stdin ⇒ `code`/`docs` — lives in [`scripts/classify-paths.sh`](../../scripts/classify-paths.sh), dependency-free and unit-tested by [`scripts/classify-paths.test.sh`](../../scripts/classify-paths.test.sh) (`make test-classify-paths`, a `verify` leaf) so the allowlists can't silently regress. The `changes` job runs the thin wrapper [`scripts/ci/classify-changes.sh`](../../scripts/ci/classify-changes.sh), which adds the CI-only policy (API file-list fetch + fail-closed: pushes, dispatches, API hiccups ⇒ `code=true`) on top. Keeping the core pure means the local git hooks can share it (`git diff --name-only | scripts/classify-paths.sh`). The `code`/`docs` outputs gate the suites and docs jobs — gate on these, never on workflow-level `paths:` filters, which would orphan the required check (invariant 1).

Expand Down
4 changes: 2 additions & 2 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -280,8 +280,8 @@ jobs:
with:
go-cache-suffix: "-e2e-cov"
# `-j` builds the prereqs (build-ts ∥ build-cover) concurrently,
# then runs the orchestrator: ClickHouse testcontainer + the cover
# binary + the SDK vitest suite.
# then runs the orchestrator: ClickHouse and Redis testcontainers +
# the cover binary + the SDK vitest suite.
- name: Build SDK dist + cover binary, run E2E suite
run: make -j "$(nproc)" test-e2e COV_DEFER=1
- name: Upload coverage fragment
Expand Down
17 changes: 11 additions & 6 deletions .testcoverage.yml
Original file line number Diff line number Diff line change
Expand Up @@ -79,9 +79,14 @@ exclude:
- ^internal/settings/
- ^cmd/wavehouse/validate\.go$
- ^cmd/wavehouse/bootstrap\.go$
# The Redis-compatible cache backend: the e2e stack runs LocalCache
# (no Redis server, and no config selects the backend yet — #613 E4),
# so the binary carries these files but e2e can never reach them; they
# pulled the e2e gate to 56.4%. The integration suite runs them against
# real servers, and the merged total still counts them.
- ^internal/cache/(redis|redis_codec|breaker|pending|metrics)\.go$
# The in-process cache backend: the e2e stack runs cache.backend=redis
# (#613), so the binary carries LocalCache and its version index but e2e
# never reaches them. The unit suite and the integration suite's main
# app (cache.backend=local) cover them; the merged total still counts them.
- ^internal/cache/(local|version_manager)\.go$
# What e2e's Redis never makes it run: the cache.redis block's rejection
# paths (boot adopts a valid fixture, as with internal/settings above)
# and the retry of invalidations the server did not take, which needs an
# outage. The unit and integration suites cover both.
- ^internal/config/cache_redis\.go$
- ^internal/cache/pending\.go$
12 changes: 6 additions & 6 deletions AGENTS.md

Large diffs are not rendered by default.

6 changes: 4 additions & 2 deletions CHANGELOG.md

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -72,7 +72,7 @@ ClickHouse is a phenomenal OLAP database, but pointing a frontend right at it le
If you're building user-facing analytics, WaveHouse is like **Supabase for ClickHouse**. Or an **open-source Tinybird** that pushes data to the frontend in real time over SSE, not just pull-based REST.

- **Ingest** — async durable WAL (embedded NATS JetStream), `200 OK` instantly, background batch-flush; schema-validated against `system.columns`; optional ID-based dedup (idempotent ingest); dead-letter queue for rows ClickHouse rejects (an unavailable ClickHouse is retried with backoff, not dead-lettered).
- **Query** — in-process Ristretto cache + `singleflight` coalescing; type-safe structured query AST; Tinybird-style named pipes (parameterized SQL endpoints).
- **Query** — result cache (in-process Ristretto, or a Redis shared by every instance) + `singleflight` coalescing; type-safe structured query AST; Tinybird-style named pipes (parameterized SQL endpoints).
- **Real-time** — native SSE push, broadcast *before* the ClickHouse flush, with JetStream gap-fill for late/reconnecting clients.
- **Security** — Hasura-style per-table, per-role column + row policies with JWT claim templating, defined in the hot-reloadable settings directory.
- **Client** — `@wavehouse/sdk`: TypeScript client with query builder, live queries, streaming, and schema codegen; one runtime dependency (an SSE frame parser, ~1.4 KB gzipped).
Expand Down
13 changes: 10 additions & 3 deletions config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -51,20 +51,27 @@ clickhouse:
password: ""
max_total_conns: 0 # ceiling on open native connections across pools; 0 = none

# Each layer's implementation, chosen at boot. Only the in-process backend
# exists for each today, and it is the default.
# Each layer's implementation, chosen at boot. The in-process backend is the
# default for each, and the only one for these three.
mq:
backend: embedded # NATS JetStream under <data_dir>/nats
dedupe:
backend: pebble # Pebble under <data_dir>/pebble
coord:
backend: local # leases (the sweeper's) held in this process

# In-process L1 cache size. The query time-bucket
# The query-result cache: local (in-process, sized by l1_max_cost) or redis
# (one Redis-compatible server shared by every instance; see the redis block
# below and the Configuration page for every key). The query time-bucket
# (query.timestamp_bucket_seconds) is a settings key.
cache:
backend: local
l1_max_cost: 67108864
# redis: # read only with backend: redis
# addrs: ["localhost:6379"] # `docker compose -f deployments/compose/dependencies.yaml --profile redis up -d`
# key_prefix: wh
# timeout: 100ms # per operation; slower is a miss, never a failed query
# The password is a secret: WH_CACHE_REDIS_PASSWORD, not this file.

# Auth has no on/off switch — the JWT middleware always runs. A request with no
# token, or an invalid/expired one, falls back to the policy default_role;
Expand Down
13 changes: 13 additions & 0 deletions deployments/compose/dependencies.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -33,5 +33,18 @@ services:
timeout: 2s
retries: 15

# Optional shared cache for trying cache.backend=redis locally, e.g. two
# host-side instances on different ports: `--profile redis`. No
# persistence, and /data on tmpfs, so it leaves no volume behind.
redis:
profiles: [redis]
# Pinned to match internal/cache's integration suite.
image: redis:8.10.2-alpine
command: ["redis-server", "--save", "", "--appendonly", "no", "--maxmemory", "256mb", "--maxmemory-policy", "allkeys-lru"]
ports:
- "6379:6379"
tmpfs:
- /data

volumes:
clickhouse-data:
8 changes: 4 additions & 4 deletions docs/src/content/docs/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -445,7 +445,7 @@ ClickHouse's inline `FORMAT` clause (e.g. `SELECT 1 FORMAT CSV` or `… FORMAT P
The proxy buffers the upstream response in memory before forwarding (no row-streaming yet), so a `SELECT *` from a large table can pin RAM on the API server. To avoid an admin OOMing themselves, responses larger than 64 MiB return 502 with a `clickhouse response exceeded N bytes` error. Narrow the query with `LIMIT`, or use a streaming client outside WaveHouse that talks to ClickHouse directly (the standard escape hatch — the same admin credentials work).
:::

This endpoint **does not cache, does not singleflight, and emits `Cache-Control: no-store`** — every request goes straight to ClickHouse, mutation or read, and downstream HTTP caches are explicitly told not to store the response. Raw SQL is an admin escape hatch with infrequent, ad-hoc traffic, so the L1/singleflight machinery would only add complexity without a real hit-rate win. Use [`POST /v1/query?table={table}`](#post-v1querytabletable--structured-query) or [`GET/POST /v1/pipes/{name}`](#getpost-v1pipesname--execute-named-pipe) for the cached read paths (dashboards, high-QPS clients, etc.) — both share an in-process L1 (Ristretto) with singleflight coalescing.
This endpoint **does not cache, does not singleflight, and emits `Cache-Control: no-store`** — every request goes straight to ClickHouse, mutation or read, and downstream HTTP caches are explicitly told not to store the response. Raw SQL is an admin escape hatch with infrequent, ad-hoc traffic, so the cache/singleflight machinery would only add complexity without a real hit-rate win. Use [`POST /v1/query?table={table}`](#post-v1querytabletable--structured-query) or [`GET/POST /v1/pipes/{name}`](#getpost-v1pipesname--execute-named-pipe) for the cached read paths (dashboards, high-QPS clients, etc.) — both go through the query cache ([`cache.backend`](/configuration#backends): in-process, or a Redis shared by every instance) with singleflight coalescing.

:::note[Admin only]
The route is mounted under `/v1/ops/*`, behind the `RequireAdmin` gate: only a caller whose JWT role equals the policy `admin_role` (`"admin"` by default) — or who presents the non-JWT [operator key](#authentication) — may use it. A tokenless request (or a valid token without a role claim) resolves to the `default_role` (not the admin role unless `default_role` is deliberately set to it — a loudly-warned dev-only setting) and is rejected with `403`; a present-but-invalid token — expired, malformed, bad signature — keeps its stashed verification error and fails loud with `401` instead. Raw SQL has no per-statement scope check (a full SQL parser would be needed to authorize predicates), so the role gate is the entire authorization story, shared with the rest of `/v1/ops/*` (see [Admin Endpoints](#admin-endpoints)). The normal surfaces for non-admin callers are `POST /v1/ingest?table={table}` for writes, `POST /v1/query?table={table}` for structured reads, and `GET/POST /v1/pipes/{name}` for pre-defined queries — none of which expose raw SQL.
Expand Down Expand Up @@ -561,7 +561,7 @@ Table, column, and alias names may contain any characters ClickHouse accepts —

**Response:**

JSON array of result rows. Top-level `DateTime`/`DateTime64` values are returned in canonical RFC 3339 UTC (`2026-06-21T04:00:00.123Z`) — `Nullable` timestamp columns included (a SQL `NULL` renders as JSON `null`), while timestamps nested inside `Array`/`Map`/`Tuple` columns are rendered in the column's declared zone, else the ClickHouse server's, as the driver returns them — byte-identical to the [SSE stream](#get-v1stream--server-sent-events-stream) for values [canonicalized at ingest](#timestamp-canonicalization) (a fail-open pass-through that ClickHouse accepted still comes back canonical here, though it streamed in the producer's spelling). The response carries an `X-Cache: HIT` or `X-Cache: MISS` header — this endpoint shares the in-process L1 (Ristretto) + singleflight machinery (unlike `/v1/ops/query`, which always hits ClickHouse), keyed by [tenant](/deployment#multi-tenant-deployments): a request is never served from, or coalesced with, another tenant's.
JSON array of result rows. Top-level `DateTime`/`DateTime64` values are returned in canonical RFC 3339 UTC (`2026-06-21T04:00:00.123Z`) — `Nullable` timestamp columns included (a SQL `NULL` renders as JSON `null`), while timestamps nested inside `Array`/`Map`/`Tuple` columns are rendered in the column's declared zone, else the ClickHouse server's, as the driver returns them — byte-identical to the [SSE stream](#get-v1stream--server-sent-events-stream) for values [canonicalized at ingest](#timestamp-canonicalization) (a fail-open pass-through that ClickHouse accepted still comes back canonical here, though it streamed in the producer's spelling). The response carries an `X-Cache: HIT` or `X-Cache: MISS` header — this endpoint shares the query cache + singleflight machinery (unlike `/v1/ops/query`, which always hits ClickHouse), keyed by [tenant](/deployment#multi-tenant-deployments): a request is never served from, or coalesced with, another tenant's.

The inbound request body is capped at 1 MiB; a body over the cap is rejected with `413`. A query AST is bounded by nature (far under 1 MiB even with a large `in`-list), and the cap blocks a single-request memory-exhaustion vector on this public endpoint. Set a tighter or higher outer limit at your [reverse proxy](/reverse-proxy#request-body-size-limits) — but it can only narrow the effective limit, not raise it past this cap.

Expand All @@ -585,7 +585,7 @@ The inbound request body is capped at 1 MiB; a body over the cap is rejected wit

### `GET/POST /v1/pipes/{name}` — Execute Named Pipe

Executes a pre-defined named query (pipe) with parameter binding. Parameters can be supplied via query string and/or JSON body. Results are cached in the shared L1 (Ristretto) with singleflight coalescing — same machinery as the structured query endpoint, keyed by [tenant](/deployment#multi-tenant-deployments) like it, and again, unlike `/v1/ops/query`.
Executes a pre-defined named query (pipe) with parameter binding. Parameters can be supplied via query string and/or JSON body. Results are cached in the query cache ([`cache.backend`](/configuration#backends): in-process, or a Redis shared by every instance) with singleflight coalescing — same machinery as the structured query endpoint, keyed by [tenant](/deployment#multi-tenant-deployments) like it, and again, unlike `/v1/ops/query`.

**Query Parameters:** Any key matching a pipe parameter name.

Expand All @@ -600,7 +600,7 @@ Executes a pre-defined named query (pipe) with parameter binding. Parameters can

**Response:**

JSON array of result rows, with `X-Cache: HIT` or `X-Cache: MISS` indicating whether the row came from the in-process L1.
JSON array of result rows, with `X-Cache: HIT` or `X-Cache: MISS` indicating whether the rows came from the query cache.

The POST parameter body is capped at 1 MiB; a body over the cap is rejected with `413` (the same 1 MiB parameter/AST-body cap as [`POST /v1/query`](#post-v1querytabletable--structured-query) — see [reverse proxy → body limits](/reverse-proxy#request-body-size-limits)). A malformed-but-within-cap body is ignored rather than rejected, since parameters may legitimately come from the query string alone.

Expand Down
Loading
Loading