I still always want to know the numbers per edge. Because there'a a reaon why these graph DBs try to hold the stuff in memory, everything else is slow as heck. We have ~30bytes per edge in our succinct indexes, which makes the 1.5gb wiki dataset clock in at ~50gb.
If I understand correctly you shard the graph and then content adress each chunk:
- How do you decide on the sharding? I'd expect finding good connected subsets with nice memory locality to be extremely difficult computationally?
- How do you canonicalise your graph. Which has also been an extremely difficult problem in RDF land for example. Although that one is at least efficiently solvable.
LLMs have no concept of focus when it comes to docs, so they spew out way more information than is necessary or helpful to a human reading the document.
Since it is a much in demand feature. Why do you think they have not done it themselves already, versus what you have done?
Curious if this is breakthrough? Or there are some negatives that prevent Neo4J from doing it also?