Its great to see scale and progress, but being closed source is HUGE DEALBREAKER.
Clickhouse is also on the right track of building some amazing opensource integrations with postgres, they have superior*[1] managed postgres looks like from their recent blog. I hope they do some OSS sharded postgres solution.
Sai from ClickHouse here, I lead the Postgres efforts at ClickHouse. Expect news from us on this soon! Many of us here are ex-Citus and have done this for Postgres before.
Most Citus workloads were 10s of TB with largest at around a few PB or so. Heap was a couple PB, back then, if I remember correctly. It is a brilliant piece of technology that supported mission critical workloads across mid/late stage startups to huge enterprises. The planner/executor are very advanced supporting a multitude of features and decade of intricate effort.
The biggest problem of Citus was migration effort, transition from single node to multi-node was not trivial. Here I’m not talking about single table use-cases, more classic relational, multi-tenant apps with 100s to 1000s of tables. This is partly expected with most sharding technologies, though.
Sharing some insights based on my multiple years of experience working with Citus!
Schema changes need locking... everywhere. Citus is no different. And they need proper design everywhere. If you mean that it requires distributed transactions, well, yes, again: expected and solved. Not even all sharding solutions support this.
> coordinator node
Not sure what the problem is here. If what you mean is that a single coordinator, even with an HA replica, can saturate, that's true, but you can add multiple "query routers" (that's our name in StackGres, see [1]).
> Maybe I am wrong, but I am yet to read stories on operating tens of TB scale workloads on citus.
For example, we have a customer that ingests some 30TB/day, and it's ramping up towards 200TB/day of ingestion. On 24 worker nodes.
From the benchmarks on 4vCPU and num_clients=4, the numbers doesn't look much different.
Reactive looks promising, doesn't look much useful in realworld for a cache.
For example, a client subscribes for something and the machines goes down, what happens to reactivity?
Is that really your reaction to C++? These days when selecting a critical infrastructure component I'm very inclined to prefer memory-safe by default, like rust or (somewhat less so) golang. Sure, C/C++ projects that have been around and under attack for years (redis, postgres, etc) are fine but they've had those years of battle testing. For a new project I really feel a lot safer if they're built in something more failsafe.
Clickhouse is also on the right track of building some amazing opensource integrations with postgres, they have superior*[1] managed postgres looks like from their recent blog. I hope they do some OSS sharded postgres solution.
[1] - https://clickhouse.com/blog/benchmarking-nvme-managed-postgr...
reply