Show HN: LatticeDB – 类似 SQLite 的图数据库
Show HN: LatticeDB – Like SQLite but for graph databases

原始链接: https://github.com/jeffhajewski/latticedb

**LatticeDB** 是一款嵌入式、单文件属性图数据库,专为高性能、本地优先(local-first)的应用程序而设计。它使开发者能够在统一的查询层和引擎中,同时执行复杂的图遍历、HNSW 向量相似度搜索以及 BM25 全文搜索。 **核心特性:** * **统一架构:** 将图、向量和文本索引整合在单个便携文件中。 * **高性能:** 基准测试显示,其节点查找速度在微秒级,100 万规模的向量搜索仅需 0.83 毫秒,在图处理和全文搜索负载方面,性能显著优于 SQLite 等传统嵌入式方案。 * **零配置:** 作为单写者、本地优先引擎运行,并具备由预写式日志(WAL)支持的 ACID 持久性。 * **开发者友好:** 支持类 Cypher 查询语言,具备持久化事件流,并提供 Python、TypeScript 和 Go 语言绑定。 **使用场景:** 非常适合关系密集型任务,如 Graph RAG(图增强检索生成)、智能体(Agent)记忆以及本地知识管理工具,尤其是在不希望引入服务器端开销的情况下。 **避免使用场景:** LatticeDB 不适用于多进程/多客户端并发写入、分布式(多机)扩展,或严格的表格式数据工作流。它是一个精简且专业的替代方案,用于取代 Neo4j 或 PostgreSQL 等服务器端数据库,而非全面替代功能完备的成熟企业级生态系统。

LatticeDB 是由 Jeff Hajewski 开发的一个新的“Show HN”项目,旨在将 SQLite 的“本地优先”便利性带入图数据库领域。出于在本地开发环境中操作图数据库时遇到的阻碍,Hajewski 开发了 LatticeDB,旨在提供一种更易于使用的替代方案。 该项目已获得 Hacker News 社区的初步好评,用户对其类 SQLite 的本地使用方式表示赞赏。讨论的焦点集中在技术可扩展性、测试以及与类似工具的比较上。具体来说,评论者提出了关于 LatticeDB 与 DuckPGQ 和 LadybugDB 等新兴技术相比如何的问题。开发者提供了明确的文档,将 LatticeDB 与另一种图数据库解决方案 Kuzu 进行了对比。总的来说,该项目满足了开发者的一个小众需求:即在不需要传统企业级架构复杂性的前提下,获得图数据库的能力。
相关文章

原文

Embedded property-graph database with native vector and full-text indexing.

LatticeDB is a single-file local database for connected, semantic, and textual data. It lets you traverse relationships, run vector similarity search, and do BM25 full-text search over the same dataset in one engine and one query layer. It is designed for relationship-heavy workloads on a single machine, with zero-config operation and an embedded single-writer model.

LatticeDB is an embedded, single-file graph database that lets local applications query the same data by relationship, semantics, and text, then consume durable graph and application events from the same file. Workloads like Graph RAG, agent memory, and local knowledge tools are examples built on those primitives, not the definition of the engine.

  • One file. Your entire database is a single portable file. No server, no configuration.
  • One query layer. Graph traversal, HNSW vector similarity, and BM25 full-text — in the same query language.
  • One event log. Durable named streams and a built-in graph changefeed share the same transaction/WAL path as graph writes.
  • Local-first. Designed for one owning process on one machine, with WAL-backed durability.
  • Fast. 0.13 μs node lookups. 0.83 ms vector search at 1M vectors with 100% recall.
-- Find chunks similar to a query, traverse to their document, then to the author
MATCH (chunk:Chunk)-[:PART_OF]->(doc:Document)-[:AUTHORED_BY]->(author:Person)
WHERE chunk.embedding <=> $query_vector < 0.3
  AND doc.content @@ "neural networks"
RETURN doc.title, chunk.text, author.name
ORDER BY chunk.embedding <=> $query_vector
LIMIT 10

CLI

curl -fsSL https://raw.githubusercontent.com/jeffhajewski/latticedb/main/dist/install.sh | bash

Python

Published wheels are expected to bundle liblattice on supported platforms. Source installs can also bundle a staged native library during wheel builds with LATTICE_BUNDLE_LIB_DIR=/path/to/lib.

TypeScript / Node.js

npm install @hajewski/latticedb

Published package tarballs are expected to bundle liblattice on supported platforms. Source checkouts can stage the native library into the package with LATTICE_BUNDLE_LIB_DIR=/path/to/lib npm run bundle:native.

Go

See bindings/go/README.md for the current cgo workflow. The default consumer path uses installed pkg-config metadata; in-repo development can use -tags repolocal against zig-out/lib. There is also a runnable graph/vector/text retrieval example in examples/go.

Recent binding-surface cleanups moved embedding helpers into dedicated modules and subpackages. See docs/client_api_migration.md for the preferred imports and current compatibility aliases.

A complete example: create a small knowledge graph with documents and authors, store embeddings, index text, then query across all three search modes.

from latticedb import Database
from latticedb.embedding import hash_embed

with Database("knowledge.db", create=True, enable_vectors=True, vector_dimensions=128) as db:

    # --- Build the graph ---
    with db.write() as txn:
        # Create authors
        alice = txn.create_node(labels=["Person"], properties={"name": "Alice", "field": "ML"})
        bob = txn.create_node(labels=["Person"], properties={"name": "Bob", "field": "Systems"})
        txn.create_edge(alice.id, bob.id, "COLLABORATES_WITH")

        # Create documents with chunks
        for title, text, author in [
            ("Attention Is All You Need", "The transformer architecture uses self-attention...", alice),
            ("Scaling Laws for LLMs", "We find that model performance scales predictably...", alice),
            ("Log-Structured Merge Trees", "LSM trees optimize write-heavy workloads...", bob),
        ]:
            doc = txn.create_node(labels=["Document"], properties={"title": title})
            chunk = txn.create_node(labels=["Chunk"], properties={"text": text})

            # Store embedding and index text
            txn.set_vector(chunk.id, "embedding", hash_embed(text, dimensions=128))
            txn.fts_index(chunk.id, text)

            txn.create_edge(chunk.id, doc.id, "PART_OF")
            txn.create_edge(doc.id, author.id, "AUTHORED_BY")

        txn.commit()

    # --- Query: vector search + text match + graph traversal ---
    results = db.query("""
        MATCH (chunk:Chunk)-[:PART_OF]->(doc:Document)-[:AUTHORED_BY]->(author:Person)
        WHERE chunk.embedding <=> $query < 0.5
        RETURN doc.title, chunk.text, author.name
        ORDER BY chunk.embedding <=> $query
        LIMIT 5
    """, parameters={"query": hash_embed("transformer attention mechanism", dimensions=128)})

    for row in results:
        print(f"{row['doc.title']} by {row['author.name']}")

    # --- Full-text search ---
    for r in db.fts_search("self-attention transformer"):
        print(f"Node {r.node_id}: score={r.score:.4f}")

    # --- Aggregations ---
    stats = db.query("""
        MATCH (doc:Document)-[:AUTHORED_BY]->(p:Person)
        RETURN p.name, count(doc) AS papers
        ORDER BY papers DESC
    """)
    for row in stats:
        print(f"{row['p.name']}: {row['papers']} papers")
import { Database } from "@hajewski/latticedb";
import { hashEmbed } from "@hajewski/latticedb/embedding";

const db = new Database("knowledge.db", {
  create: true,
  enableVectors: true,
  vectorDimensions: 128,
});
await db.open();

// Build a graph
await db.write(async (txn) => {
  const alice = await txn.createNode({
    labels: ["Person"],
    properties: { name: "Alice", field: "ML" },
  });
  const doc = await txn.createNode({
    labels: ["Document"],
    properties: { title: "Attention Is All You Need" },
  });
  const chunk = await txn.createNode({
    labels: ["Chunk"],
    properties: { text: "The transformer architecture uses self-attention..." },
  });

  await txn.setVector(chunk.id, "embedding", hashEmbed("transformer self-attention", 128));
  await txn.ftsIndex(chunk.id, "The transformer architecture uses self-attention...");

  await txn.createEdge(chunk.id, doc.id, "PART_OF");
  await txn.createEdge(doc.id, alice.id, "AUTHORED_BY");
});

// Query across vector search + graph traversal
const results = await db.query(
  `MATCH (chunk:Chunk)-[:PART_OF]->(doc:Document)-[:AUTHORED_BY]->(author:Person)
   WHERE chunk.embedding <=> $query < 0.5
   RETURN doc.title, chunk.text, author.name
   ORDER BY chunk.embedding <=> $query
   LIMIT 5`,
  { query: hashEmbed("attention mechanism", 128) }
);

for (const row of results.rows) {
  console.log(`${row["doc.title"]} by ${row["author.name"]}`);
}

await db.close();
db, err := latticedb.Open("knowledge.db", latticedb.OpenOptions{
    Create: true,
    EnableVectors: true,
    VectorDimensions: 128,
})
if err != nil {
    log.Fatal(err)
}
defer db.Close()

err = db.Update(func(tx *latticedb.Tx) error {
    node, err := tx.CreateNode(latticedb.CreateNodeOptions{
        Labels: []string{"Chunk"},
        Properties: map[string]latticedb.Value{"text": "The transformer architecture uses self-attention..."},
    })
    if err != nil {
        return err
    }
    if err := tx.SetVector(node.ID, "embedding", []float32{1, 0, 0, 0}); err != nil {
        return err
    }
    return tx.FTSIndex(node.ID, "The transformer architecture uses self-attention...")
})
if err != nil {
    log.Fatal(err)
}

Benchmarked on Apple M1, single-threaded, with auto-scaled buffer pool. Run zig build benchmark to reproduce. For the repeated-term FTS indexing workload that previously exposed quadratic append behavior, run zig build fts-benchmark.

Operation Latency Throughput Target Status
Node lookup 0.13 μs 7.9M ops/sec < 1 μs PASS
Node creation 0.65 μs 1.5M ops/sec
Edge traversal 9 μs 111K ops/sec
Full-text search (100 docs) 19 μs 53K ops/sec
10-NN vector search (1M vectors) 0.83 ms 1.2K ops/sec < 10 ms @ 1M PASS

Vector Search (HNSW) at Scale

128-dimensional cosine vectors, M=16, ef_construction=200, ef_search=64, k=10. Run zig build vector-benchmark to reproduce.

Scale Mean Latency P99 Latency Recall@10 Memory
1,000 65 μs 70 μs 100% 1 MB
10,000 174 μs 695 μs 99% 10 MB
100,000 438 μs 1.2 ms 99% 101 MB
1,000,000 832 μs 1.8 ms 100% 1,040 MB

Search latency scales sub-linearly (O(log N)) with 99–100% recall@10. Uses heuristic neighbor selection (HNSW paper Algorithm 4) for diverse graph connectivity, connection page packing for ~4.5x memory reduction, and pre-normalized dot product for fast cosine distance.

ef_search Sensitivity (1M vectors)

ef_search Mean Latency Recall@10
16 506 μs 57%
32 1.9 ms 79%
64 990 μs 100%
128 3.2 ms 100%
256 11.6 ms 100%
System Latency Type Source
LatticeDB 0.13 μs Embedded zig build benchmark
RocksDB (in-memory) 0.14 μs Embedded RocksDB wiki
SQLite (in-memory) ~0.2 μs Embedded Turso blog
SQLite (WAL, disk) 3 μs (p90) Embedded marending.dev
Neo4j 28 ms (p99) Server Memgraph comparison

LatticeDB's B+Tree achieves sub-microsecond cached lookups, matching RocksDB in-memory and outperforming SQLite on disk by 23x.

LatticeDB at 1M achieves 0.83 ms mean with 100% recall@10 — faster than FAISS single-threaded HNSW and competitive with Weaviate and Qdrant server-based systems (which add network overhead in practice).

System 2-hop (100K nodes) Type Source
LatticeDB 39 μs Embedded zig build sqlite-benchmark
SQLite (recursive CTE) 548 μs Embedded zig build sqlite-benchmark
Kuzu 19 ms Embedded The Data Quarry
Neo4j 10 ms (1M nodes) Server Neo4j blog

LatticeDB vs SQLite — Social network graph with power-law degree distribution, adjacency cache pre-warmed:

Small Scale (10K nodes, 50K edges)

Workload LatticeDB SQLite Speedup
1-hop traversal 560 ns 13.0 μs 23x
2-hop traversal 3.0 μs 37.5 μs 13x
3-hop traversal 19.1 μs 178.5 μs 9x
Variable path (1..5) 82.4 μs 4.3 ms 52x

Medium Scale (100K nodes, 500K edges)

Workload LatticeDB SQLite Speedup
1-hop traversal 8.0 μs 290.0 μs 36x
2-hop traversal 38.7 μs 548.3 μs 14x
3-hop traversal 197.3 μs 1.2 ms 6x
Variable path (1..5) 134.4 μs 10.1 ms 75x

Depth-Limited Traversal (10K nodes, 50K edges)

Depth LatticeDB SQLite Speedup
10 311 μs 121 ms 390x
15 380 μs 271 ms 713x
25 318 μs 587 ms 1,848x
50 500 μs 1.4 s 2,819x

LatticeDB uses BFS with adjacency cache and bitset visited tracking. SQLite uses a recursive CTE with UNION deduplication. Both compute identical reachable node sets (~8K nodes). The gap widens at deeper depths as SQLite's CTE overhead grows with each recursion level. Run zig build graph-benchmark -- --quick to reproduce.

System Search Latency Type Source
LatticeDB 19 μs Embedded zig build benchmark
SQLite FTS5 < 6 ms Embedded SQLite Cloud
Elasticsearch 1–10 ms Server Various
Tantivy 10–100 μs Library Various

LatticeDB's inverted index with BM25 scoring is ~300x faster than SQLite FTS5 and competitive with Tantivy (a dedicated Rust search library).

Graph

  • Nodes and edges with labels and arbitrary properties
  • Durable explicit equality indexes for scoped node and edge properties
  • Multi-hop traversal, variable-length paths (*1..3)
  • ACID transactions with commit/rollback and crash recovery
  • MERGE, WITH, UNWIND, aggregations (count, sum, avg, min, max, collect)

Vector Search

  • HNSW approximate nearest neighbor with configurable M, ef
  • Built-in hash embeddings or HTTP client for Ollama/OpenAI
  • Bulk vector node insertion for fast ingestion

Full-Text Search

  • BM25-ranked inverted index with tokenization and stemming
  • Fuzzy search with configurable Levenshtein distance

Cypher Query Language

  • MATCH, WHERE, RETURN, CREATE, DELETE, SET, REMOVE
  • ORDER BY, LIMIT, SKIP, DETACH DELETE
  • Vector distance operator: <=>
  • Full-text search operator: @@
  • Parameters: $name

Operations

  • Single-file storage with write-ahead log for crash recovery
  • Durable named streams with explicit consumer offsets, manual trim, and graph changefeeds
  • Online freelist reuse plus lattice compact for safe physical tail reclamation
  • Zero configuration — open a file and start working
  • Embedded single-writer model for local applications
  • Clean C API; Python, TypeScript, and Go bindings wrap it
  • Connected local data — Notes, documents, catalogs, citation graphs, and entity graphs
  • Graph plus retrieval — Relationship traversal, semantic search, and lexical search over the same dataset
  • Local knowledge tools — Embedded apps that need graph structure without running a separate server
  • Agent memory and RAG pipelines — One example class of workload built on the graph/vector/text substrate
  • Local development — Lightweight alternative to Neo4j or Weaviate for prototyping on one machine

When to Use Something Else

LatticeDB is fast, but speed is not the only thing that matters. Here are cases where a different tool is the better choice.

You need multiple applications writing to the same database at the same time. LatticeDB is embedded with a single-writer model. One process opens the file and owns it. If you need many clients connecting over a network, use Neo4j, PostgreSQL, or another client-server database.

Your data is fundamentally tabular. If your data fits naturally into rows and columns — sales records, user accounts, time series — a relational database like SQLite or PostgreSQL will be simpler and just as fast. Graph databases shine when relationships between records are the point, not an afterthought.

You need to scale beyond a single machine. LatticeDB stores everything in one file on one machine. If you need sharding, replication, or distributed queries across billions of nodes, look at Neo4j cluster, Dgraph, or a managed service like Neptune.

You need the full Cypher language. LatticeDB supports most of Cypher but not all of it. Features like OPTIONAL MATCH and CALL procedures are not yet implemented. If your queries depend on these, Neo4j is the complete implementation.

You need mature tooling and ecosystem. Neo4j has visualization tools, admin dashboards, monitoring, drivers in every language, and years of community resources. PostgreSQL has decades of tooling. LatticeDB is new and lean — which is a strength for embedding, but a weakness if you need a rich operational ecosystem around your database.

Written in Zig. No dependencies.

git clone https://github.com/jeffhajewski/latticedb.git
cd latticedb
zig build                  # build everything
zig build test             # run tests
zig build -Doptimize=ReleaseFast   # optimized build

MIT

联系我们 contact @ memedata.com