Hacker Newsnew | past | comments | ask | show | jobs | submit | weekendcode's commentslogin

Its great to see scale and progress, but being closed source is HUGE DEALBREAKER.

Clickhouse is also on the right track of building some amazing opensource integrations with postgres, they have superior*[1] managed postgres looks like from their recent blog. I hope they do some OSS sharded postgres solution.

[1] - https://clickhouse.com/blog/benchmarking-nvme-managed-postgr...


Sai from ClickHouse here, I lead the Postgres efforts at ClickHouse. Expect news from us on this soon! Many of us here are ex-Citus and have done this for Postgres before.

If you want an OSS sharded Postgres, just use Citus. Feature-full, proven, mature, boring.

its not feature-full. like schema changes locking, coordinator node..

Maybe I am wrong, but I am yet to read stories on operating tens of TB scale workloads on citus.


Most Citus workloads were 10s of TB with largest at around a few PB or so. Heap was a couple PB, back then, if I remember correctly. It is a brilliant piece of technology that supported mission critical workloads across mid/late stage startups to huge enterprises. The planner/executor are very advanced supporting a multitude of features and decade of intricate effort.

The biggest problem of Citus was migration effort, transition from single node to multi-node was not trivial. Here I’m not talking about single table use-cases, more classic relational, multi-tenant apps with 100s to 1000s of tables. This is partly expected with most sharding technologies, though.

Sharing some insights based on my multiple years of experience working with Citus!

Here are few customer use-cases I could found:

https://docs.citusdata.com/en/v10.0/get_started/what_is_citu...

https://info.citusdata.com/rs/235-CNE-301/images/Citus_Data_...?

https://www.youtube.com/watch?v=F6df3HV6kP0


> schema changes locking

Schema changes need locking... everywhere. Citus is no different. And they need proper design everywhere. If you mean that it requires distributed transactions, well, yes, again: expected and solved. Not even all sharding solutions support this.

> coordinator node

Not sure what the problem is here. If what you mean is that a single coordinator, even with an HA replica, can saturate, that's true, but you can add multiple "query routers" (that's our name in StackGres, see [1]).

> Maybe I am wrong, but I am yet to read stories on operating tens of TB scale workloads on citus.

For example, we have a customer that ingests some 30TB/day, and it's ramping up towards 200TB/day of ingestion. On 24 worker nodes.

[1]: https://stackgres.io/doc/latest/administration/sharded-clust...


Interesting, nice work on the small details!

Do you happen to have a technique post on what goes behind this tech?


Glad you liked it. Talking about the Technical post, Not yet but is on the cards. Will definitely work on it and post it here, thanks again.


ofcourse it has it rewritten in rust.


From the benchmarks on 4vCPU and num_clients=4, the numbers doesn't look much different.

Reactive looks promising, doesn't look much useful in realworld for a cache. For example, a client subscribes for something and the machines goes down, what happens to reactivity?


I guess yes.


Built in C++, awesome :)


Is that really your reaction to C++? These days when selecting a critical infrastructure component I'm very inclined to prefer memory-safe by default, like rust or (somewhat less so) golang. Sure, C/C++ projects that have been around and under attack for years (redis, postgres, etc) are fine but they've had those years of battle testing. For a new project I really feel a lot safer if they're built in something more failsafe.


Can you show me where modern C++ is unsafe and why it matters in this specific context?



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: