TollMeshCache vs Redis

An honest comparison, current as of this project’s actual tested state — not aspirational.


What’s Actually Verified Working

Fourteen capabilities are wired end-to-end (Go backend → HTTP API → all 7 SDKs) and have been run against a live server as part of building them:

Feature TollMeshCache Redis
Rate limiting Yes (CRDT GCounter) Yes (requires INCR+EXPIRE or a module)
Replay protection Yes (CRDT GSet) Yes (requires SETNX+EXPIRE)
Key-value cache with TTL Yes Yes
Job queues Yes (priority, retry, dead-letter) Not built in (commonly layered on Streams or Lists)
Sorted sets Yes (skip list, CRDT conflict resolution) Yes (skip list)
Streams with consumer groups Yes Yes
Pub/Sub Yes (poll-based delivery, not a persistent connection) Yes (PUBLISH/SUBSCRIBE over RESP)
Transactions Yes (queue ops, atomic commit/rollback) Yes (MULTI/EXEC, no rollback on runtime errors)
Persistence Yes (WAL + snapshot, checksummed) Yes (RDB + AOF, 15+ years battle-tested)
Scripting Yes — real arbitrary Go code, compiled by TinyGo to WASI WASM, executed in a sandboxed wazero runtime Yes (Lua via EVAL)
Search (BM25 + vector, hybrid) Yes No (requires RediSearch module)
Ranking Yes No (build it yourself)
Metrics (JSON + Prometheus) Yes Via INFO command or exporter

For all fourteen, correctness is backed by real tests (Go unit tests for the backend, plus live HTTP integration tests run against every SDK, not just “it compiles”).

Authentication is real too, as of a security review this session — worth calling out on its own, since it wasn’t true for most of this project’s history. Every SDK has sent an X-API-Key header since it was written, but the server never checked it: every request succeeded regardless of whether a key was sent or correct. Fixed — see API Reference: Authentication — and while fixing it, found and fixed four SDKs that couldn’t actually have authenticated even if the server had been checking: Rust and Java accepted an api_key config value and silently never sent it anywhere, and PHP only sent it on POST requests, never GET (so cache_get, health, and every other read-only call went out unauthenticated). All 7 SDKs are now live-verified end-to-end against a real server with an API key enabled.


Scripting: a genuinely different design from Redis, on purpose

Redis scripting is Lua via EVAL/EVALSHA. This project deliberately does not depend on Redis or a Redis-derived component anywhere, including for scripting, so instead of embedding a Lua VM, a script here is Go source code: compiled server-side by the TinyGo toolchain to a WASI WebAssembly module, then executed in a sandboxed wazero runtime (pure Go, no cgo), with a hard execution timeout and memory limit enforced per call.

This mirrors the shape of Redis’s SCRIPT LOAD + EVALSHA split — compile once (slow, real seconds, since it invokes an external compiler), execute many times (fast, single-digit milliseconds, since it reuses the already-compiled module) — but the execution surface is genuinely arbitrary Go, not a restricted scripting language, sandboxed by WASM isolation rather than a Lua interpreter’s built-in restrictions. An infinite loop or any script exceeding its timeout is force-terminated without affecting the server process, other scripts, or other cached compiled modules.

For cases that don’t need arbitrary code — composing the server’s own built-in operations (get, set, zadd, enqueue, …) into a multi-step sequence — there’s also a separate, simpler Pipeline primitive with no code-execution surface at all.


Performance

A first real, reproducible benchmark pass now exists (api/http_bench_test.go, persistence/wal_bench_test.go; run with go test ./api/... -bench BenchmarkHTTP -benchtime=3000x -run '^$' and similarly for ./persistence/...) — measured on one Apple M2 laptop, not production hardware, and not a substitute for benchmarking your own deployment target. Any other throughput or latency numbers you might see elsewhere in this project’s history that aren’t backed by a runnable benchmark in the repo should still not be trusted.

What the numbers say, and what can be said honestly beyond them:

  • A real HTTP round trip (loopback TCP, not an in-process handler call) to /consume — one of the simplest writes in the system — takes ~45μs sequential, ~15.6μs amortized under concurrent load (~64k ops/sec aggregate on 8 cores). /cache/set + /cache/get back to back takes ~88-93μs. This is a real floor under latency compared to Redis’s RESP protocol, which keeps a persistent connection open and doesn’t pay TCP/HTTP overhead per command; TollMeshCache has no persistent-connection or pipelining protocol yet.
  • SortedSet.Rank/RevRank are now O(log n), matching Redis’s ZRANK (Insert, Delete, and range queries were already O(log n)). This required adding per-pointer span tracking to the skip list, which also surfaced and fixed a real correctness bug: Delete previously navigated by member name alone even though the list is ordered by (score, member), so a member whose name sorted differently than its score position could silently fail to be removed (confirmed live, and now covered by a randomized regression test — see sortedset/skiplist_test.go).
  • WASM script execution, once compiled, is fast (single-digit milliseconds observed live across all 7 SDKs) because the compiled module is cached and reused; compilation itself takes real seconds (TinyGo invokes an actual Go-to-WASM compiler process) and should be treated the same way you’d treat SCRIPT LOAD — infrequent, not on the hot path.
  • Every successful Set/Consume/Seen now logs to the WAL (PersistenceEngine.LogOperation, ~4.4μs sequential, ~3.7μs under concurrent writers), and a fresh process automatically recovers full state on startup by loading the latest snapshot and replaying every WAL entry after it — verified live by hard-killing a running server (SIGKILL, no graceful shutdown) and confirming a new process on the same data directory recovered cache, replay-protection, and exact rate-limiter counts correctly. This was not true until recently: for most of this project’s history “write-ahead log” was aspirational — the write path never called it, so only an explicit create_snapshot protected anything, and nothing replayed automatically on restart. (The package used to also contain a second, entirely unused persistence implementation — WriteAheadLog/SnapshotManager/RecoveryManager — with zero references outside its own tests; it’s been removed rather than left as confusing dead code.) Recovery time and write-path overhead at real (non-benchmark) scale still haven’t been measured; Redis’s RDB/AOF persistence remains battle-tested over 15+ years in production at massive scale in a way this cannot yet claim.
  • cmd/loadtest drives a real, mixed, sustained concurrent workload against a separately-running server process (unlike the in-process Go benchmarks above) and immediately found a real bug: JobQueue.sortPendingJobs was a hand-rolled O(n²) double loop run on every single Enqueue, so a queue accumulating pending jobs faster than they’re claimed (an ordinary situation, not a pathological one) degraded quadratically — measured live at enqueue p50 392ms / p99 1.68s, three orders of magnitude worse than every other endpoint’s tens-of-microseconds under the identical load. Fixed with sort.Slice (O(n log n), identical resulting order); re-running the same load test afterward: overall throughput went from ~1,351 req/s to ~24,461 req/s (18x), and enqueue’s p99 dropped to 44ms.

Architecture

TollMeshCache: peer-to-peer, CRDT-based. Each node holds full state; nodes gossip over the same HTTP API every SDK uses and merge incoming state via each primitive’s real CRDT merge (GCounter/GSet, and cache’s per-key LWW-register). No coordinator, no leader election. Peer health is now real too: PeerManager (previously built but never wired to anything, its health check literally a “Simulate” placeholder) tracks every peer via both gossip’s own round-trip results and an independent periodic /health probe, backing genuine /livez (process-alive) and /readyz (not-isolated-from-the-cluster) endpoints — verified live by killing and restarting a node and watching the survivor’s readiness flip and recover accordingly. Transport encryption is now opt-in too: -tls-cert/-tls-key/-tls-ca flags turn on TLS for the HTTP API and for gossip/health-check traffic between nodes, verified against a shared CA rather than trusted blindly — previously everything, including gossip’s node-to-node state transfer, was plaintext. When all three flags are set, TLS is mutual: a node’s server now requires and verifies every caller’s client certificate too, not just the other direction, closing the gap where completing a TLS handshake at all (without proving cluster membership via a certificate) was enough to reach a node’s HTTP surface. Certificates also hot-reload now (coordination.CertReloader), watching the cert/key files’ mtimes and re-parsing on change – verified live by rotating a running node’s certificate on disk and confirming it presented the new one within seconds, no restart. A broader security review pass (the closest honest proxy for external audit available here) found one real newly-introduced issue: /debug/pprof/cmdline remotely exposed this process’s command-line arguments – including -cluster-secret if passed as a flag – to anyone holding just the API key, a materially weaker credential than the cluster secret; fixed by moving /debug/pprof/* and the related /peers/health to the cluster-secret tier. A 10-minute, 2.3M-request sustained load test (cmd/loadtest, unlike the earlier 20-second run) also found two more real bugs of the same shape as before — a configured value that was stored and even reported via stats but never actually enforced: JobQueue.maxAge (completed/failed jobs never evicted from memory) and PersistenceEngine.snapshotInterval (no periodic auto-snapshot ever ran, so the WAL grew to ~213MB with zero snapshots taken over the test). Both fixed; goroutine count stayed flat the entire 10 minutes, confirming no leak there. A 3-node concurrent load test (all three nodes hammered simultaneously with real gossip converging live) found zero errors across ~2M total requests and exact metric convergence across every node. Job Queues’ and Transactions’ documented cross-node race window (waiting on periodic gossip alone) also got a real, bounded mitigation: an out-of-band, debounced gossip push after a claim or commit, cutting the typical window from ~5s to ~45ms measured live — a large reduction, not an elimination. Gossip’s real scalability ceiling — it transfers the entire replicated state every round, not a delta, so cost scales with total data volume rather than with what actually changed — remains open (fixing it means per-key delta tracking across all thirteen replicated primitives, out of scope for one pass), but the actual bytes on the wire are now meaningfully smaller: /internal/state responses gzip-compress automatically when the requester supports it (every client here does, with zero code changes needed on that side), measured live at ~89% smaller against a populated real state. A second file that looked like it might already solve the delta-sync problem, coordination/state_sync.go’s Merkle-tree diffing, turned out not to work at all as a diff mechanism (non-deterministic tree construction, no tree-comparison method anywhere) and covered less than the real merge path besides — removed rather than left as misleading dead code. This is a real, meaningful structural difference from Redis, and — unlike a lot of what’s in this document’s history — it’s now genuinely working, not just unit-tested in isolation: verified live against real, separately launched OS processes, including concurrent writes to the same key on different nodes converging correctly and correct counter aggregation across the cluster. Thirteen primitives replicate today: the original three (rate limiting, replay protection, cache) plus, as of this session, Sorted Sets, Streams, Pipelines, Search, Job Queues, Pub/Sub, Transactions, WASM Scripting, (partially) Metrics, and named Ranking configs. Sorted Sets was straightforward — SortedSet.Merge already had a real CRDT conflict resolution built for local testing, just not wired to gossip yet. Streams needed a real fix first: entry IDs weren’t actually globally unique across nodes (a plain per-node sequence counter, easy to collide across two nodes writing in the same millisecond), which would have silently corrupted a union-merge — fixed by putting the node into the ID. Pipelines was close to Cache’s shape (a named registry, LWW per entry) and mostly reused that pattern, with one open gap: pipeline deletion doesn’t replicate (no tombstone yet). Search was the same shape as Pipelines (a per-document LWW register), but needed a real pre-existing bug fixed first — re-indexing a document under an ID that was already indexed double-counted every BM25 statistic instead of replacing the old contribution — and shares the same open gap: document deletion doesn’t replicate. Job Queues replicates each queue’s job log the same LWW way, but its open gap is bigger than a stale-read: merging state doesn’t provide exclusive claims across nodes, so two nodes can each claim the same job before gossip converges, making processing at-least-once across the cluster rather than exactly-once. Pub/Sub needed the same entry-ID uniqueness fix as Streams, and converges each topic’s message history (a set union by message ID) across nodes — but deliberately does not attempt to push merged messages into a live Subscriber channel, since that channel is in-process memory with no cross-node serialization story; Subscribe/Poll for a given subscriber must still land on the same node. Transactions needed a real durability fix first — a committed transaction’s Set effects were never logged to the WAL, unlike a plain Set, so they were silently lost on restart — and shares Job Queues’ shape of gap: transaction metadata converges eventually, but not cross-node atomicity, so a transaction’s begin/add-operation/commit calls must land on the same node. WASM Scripting is structurally different from all the others: there’s no compiled artifact to gossip, so only a script’s Go source and version metadata replicate, and adopting a peer’s version means actually invoking the TinyGo compiler locally (real seconds, not a cheap struct swap) — a cluster mixing nodes with and without TinyGo simply never converges scripts onto the nodes lacking it. Its merge is also insert-only, unlike every other feature’s LWW-overwrite merge: a security review found that letting a peer’s “newer” claim overwrite an existing script would let a node holding only the cluster secret hijack a script name that was only ever supposed to be set via the separately-credentialed, API-key-gated compile path — since merging a script, unlike merging inert data, runs the compiler as a direct side effect. Fixed so gossip can only introduce a genuinely new script name, never replace one that already exists locally. Metrics is a partial, different-shaped case: its 13 monotonic counters replicate as a real grow-only-counter CRDT (max per node per metric, the same rule as rate limiting’s GCounter), giving a genuine cluster-wide total via a new GetClusterMetrics//metrics/cluster, but latency percentiles are deliberately excluded — merging independently-computed p50/p99 values from different nodes’ sample sets doesn’t produce a meaningful cluster percentile. Named Ranking configs are a genuinely new feature, not a fix: Rank itself is stateless (a fresh Ranker is built from caller-supplied arguments every call, nothing persists), so there was nothing to replicate there – what’s new is ranking.Registry, a named (strategy, boosts) pair registered once and referenced by name later, the same shape and the same open gap as Pipelines (deletion doesn’t replicate). Persistence is the one feature group that was never a candidate for gossip replication in the same sense as the others: its WAL/snapshot files are a per-node durability mechanism, not data with its own CRDT. Persistence’s own mechanism did have a real, separate gap, though: snapshot/restore only ever covered the original three primitives, so a lone node (or an entire cluster restarting at once) silently lost all nine newer feature groups (ranking configs included, once they existed) on every restart – now closed by feeding each one back through its own already-built MergeSnapshot, the same operation gossip already uses. See Architecture: Gossip and state sync for the details of all this, including a subtle cache durability bug a live crash test caught and a Metrics regression bug the persistence fix’s own test caught.

Redis: single-writer master with optional replicas, or Redis Cluster for horizontal scale via hash-slot sharding. Mature, well-understood failure modes, an enormous ecosystem (Sentinel, Cluster, modules, every major language’s client libraries), and production deployments at massive scale.


Honest Verdict

Redis is not being outperformed or outclassed here, and no credible claim to that effect should be made. Redis is a mature, extremely fast, heavily battle-tested system with a huge feature surface and 15+ years of production hardening. This project does not have verified multi-node operation under real load, a formal performance benchmark, or anywhere near Redis’s operational track record.

What TollMeshCache offers that’s genuinely different: a peer-to-peer CRDT model with no central coordinator, real (not Lua-derived) arbitrary-code scripting via TinyGo/WASM sandboxing, and identical client APIs across 7 languages, for fourteen capabilities that are now wired end-to-end and tested rather than aspirational — including persistence, pub/sub, and transactions, which were previously listed here as “exists as code but isn’t usable yet.”

Use TollMeshCache if the no-central-coordinator architecture or the WASM-sandboxed scripting model specifically matters to your use case, and you’ve verified the supported features cover what you need. Use Redis for anything requiring a proven production track record, multi-node operation at scale, or raw throughput.