Skip to content

Latest commit

 

History

445 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

unitdb GoDoc Go Report Card

Unitdb is blazing fast specialized time-series database for microservices, IoT, and realtime internet connected devices. As Unitdb satisfy the requirements for low latency and binary messaging, it is a perfect time-series database for applications such as internet of things and internet connected devices. The Unitdb Server uses uTP (unit Transport Protocol) for the Client Server messaging. Read uTP Specification.

Don't forget to ⭐ this repo if you like Unitdb!

About unitdb

Key characteristics

  • 100% Go
  • Can store larger-than-memory data sets
  • Optimized for fast lookups and writes
  • Supports writing billions of records per hour
  • Supports opening database with immutable flag
  • Supports database encryption
  • Supports time-to-live on message entries
  • Supports writing to wildcard topics
  • Data is safely written to disk with accuracy and high performant block sync technique

Quick Start

To build Unitdb from source code use go get command.

go get github.com/unit-io/unitdb

Usage

Detailed API documentation is available using the go.dev service.

Make use of the client by importing it in your Go client source code. For example,

import "github.com/unit-io/unitdb"

Unitdb supports Get, Put, Delete operations. It also supports encryption, batch operations, and writing to wildcard topics. See usage guide.

Samples are available in the examples directory for reference.

Running the server

The server signs client IDs and topic keys with a key only it knows, and refuses to start without one. Generate a key and set it as encryption_config's key in unitdb.conf, or in the UNITDB_ENCRYPTION_KEY environment variable:

> export UNITDB_ENCRYPTION_KEY=$(openssl rand -base64 24)

Up to v0.3 the sample unitdb.conf shipped with a key, which the server now refuses: it is public, so anyone could sign client IDs and topic keys with it. A deployment that ran with it needs a new key, and its clients new client IDs and topic keys.

Keyring and key rotation

The single key is a keyring of one key. To rotate keys, give the server a keyring instead, in the UNITDB_KEYRING environment variable or in a file named by encryption_config's keyring_file: a JSON list of keys, each with an id from 0 to 255, 32 random bytes in base64 (openssl rand -base64 32), and a use, issue for the one key the server issues with or read for keys it only reads with. The keyring takes the place of the single key, and every node of a cluster needs the same one:

> export UNITDB_KEYRING='[{"id": 1, "key": "<openssl rand -base64 32>", "use": "issue"},
                          {"id": 0, "key": "<the old key, base64>", "use": "read"}]'

The single key is key 0, used as its 32 characters are: in a keyring it is printf %s "$UNITDB_ENCRYPTION_KEY" | base64. To rotate:

  1. Add a new key as the issue key, and keep the old one as a read key. Restart every node with the keyring. Client IDs and topic keys of the old key keep working; new ones are issued with the new key.
  2. Hand out new client IDs and topic keys. A client that connects with a client ID of the old key is sent the same ID sealed with the new key on unitdb/clientid/ (see below); topic keys are requested again with unitdb/keygen. Service IDs are sealed again with mintid -from <id>.
  3. Remove the old key. What it issued is refused from then on.

The server derives a subkey of each key for each use (HKDF-SHA256): one seals client IDs, another signs topic keys, a third seals stored records (below).

Encryption at rest

Every record the server stores is sealed before the storage engine sees it: messages and their replicas, hints for other nodes, the ids of replicated messages, the topic index, the security state, sessions and their logs, and subscriptions. encrypt_at_rest is on unless unitdb.conf sets it to false: since v0.7.0, where v0.6.0 had it off unless set to true. A store written by v0.6.0 with it off opens as it is: its plain records are read, and what the server stores from then on is sealed. A deployment that wants v0.6.0's behaviour sets "encrypt_at_rest": false before it upgrades.

  • A record is sealed with XChaCha20-Poly1305 under a random 24-byte nonce, with the store subkey of the keyring's issue key, and with its contract (or its key, for sessions and logs) as associated data, so a record moved elsewhere in the store fails to open. A sealed record is magic (4) | key id (1) | nonce (24) | sealed record | tag (16): 45 bytes more than the record.
  • Records are opened with the key they name, whether encrypt_at_rest is on or off. During a rotation, records sealed with the old key open with it as a read key. Once a key is removed from the keyring, the records it sealed are refused: they are skipped, and logged as sealed with a key that is not in the keyring, rather than read as garbage. Messages expire, but sessions and subscriptions are only sealed again when they are written again, so keep an old key as a read key while the store may hold records it sealed.
  • Turning it on for an existing store leaves the records already stored as they are, readable; new ones are sealed. Turning it off again stores new records plain, and keeps reading the sealed ones while their key is in the keyring. The server doesn't seal or unseal what it stored before.
  • Topics, keys and ids are not sealed, nor are record sizes and times: only what is stored under them. Records stored before it was turned on stay plain until they expire or are written again.
  • Don't roll a server back to a version without encrypt_at_rest (v0.5.0 or before) once it has sealed records: such a version reads them as they are, sealed. With it on by default, a v0.7.0 node seals from its first start: a rollback to v0.6.0 reads its sealed records, with the same keyring.
  • In a cluster, nodes send each other records opened, and each node seals what it stores as it is set to, so nodes can be turned on one at a time, and a cluster can mix nodes with it on, off, or of an earlier version. Every node needs the same keyring already. Traffic between nodes is not encrypted by this.
  • Cost: sealing a record takes about 0.7 µs for 64 bytes and 1.8 µs for 1 KB on one core of an Apple M-series CPU, opening it a little less, and large records go at about 0.9 GB/s (go test ./server/internal/store -bench Seal).

The server doesn't use the storage engine's own encryption (unitdb.WithEncryption): its nonce is derived from the plaintext, and repeats at message volumes.

Client IDs and topic keys

The server issues and takes v2 client IDs and topic keys only. Both are opaque strings to clients:

  • A v2 client ID is 94 characters of base64url (A-Z a-z 0-9 - _). It is sealed with XChaCha20-Poly1305 under a random nonce, and holds the key id that sealed it, the contract, the permissions, a random uuid, and when it was issued and expires. client_id_ttl sets how long the IDs the server issues last (unitdb/clientid), and primary_id_ttl primary ones; they never expire by default. A client that connects with an ID of a key being retired, or past 80% of its ID's lifetime, is sent a new ID on unitdb/clientid/: the same ID, so the same contract and sessions. An expired ID is refused with return code 0x02, and is not replaced: its client needs a new one from its primary client.
  • A v2 topic key is 48 characters of base64url. Its 128-bit tag covers the whole topic and the contract, so it opens exactly the topic it was issued for (a key for ... reads every topic of the contract); it holds the key id, a uuid, and when it was issued and expires. A keygen request's ttl, such as {"topic": "teams.alpha", "type": "rw", "ttl": "24h"}, sets how long the key lasts; topic_key_ttl is the default, and keys never expire without either.

Since v0.7.0, v1 client IDs and topic keys, and unsigned keys, are refused. A v1 client ID (52 characters of base32) is refused at CONNECT with return code 0x02, and its client is not sent a new ID, which would be of another contract; on unitdb/service it vouches for nothing (status 403). A v1 signed key (26 characters) or an unsigned one (13) is refused with status 401, and the error notice on unitdb/error/ says they are no longer accepted. accept_unsigned_keys is gone: a config that still sets it to true stops the server at start, saying why, and one that sets it to false starts with a warning to remove it. v0.6.0 renews a client's v1 ID as v2 when it connects, and issues v2 keys; mintid -from <v1 id> seals a v1 ID again as v2, the same ID, for IDs kept in configs. See upgrading to v0.7.0.

Since v2 client IDs carry a uuid, two secondary IDs of a contract issued in the same second are different IDs with sessions of their own; v1 ones were the same ID. An ID sealed again from a v1 one has no uuid: it keeps the sessions of the v1 ID.

Revocation

A contract's primary client (a service ID from mintid -service is one) revokes its contract's client IDs and topic keys by publishing to unitdb/revoke:

  • {"uuid": "<uuid>"} revokes the v2 client ID or topic key with that uuid, which unitdb/clientid and unitdb/keygen answer with ("uuid", in decimal); "until": <unix seconds> ends the revocation then, for a key or ID that expires anyway.
  • {"all": true} revokes everything the contract issued before now, IDs sealed again from v1 ones included. Issue times are whole seconds, so what is issued in the same second as the request is still taken, and the nodes' clocks should agree.

The server answers {"status": 200}, 403 to a client that isn't primary (a connection a service vouched for included), and 400 to a request with nothing to revoke. A revoked or not-before ID is refused at CONNECT with return code 0x02, without a new ID, and on unitdb/service; a revoked key is refused with status 401. Connections and subscriptions already open stay until they reconnect or subscribe again. An ID sealed again from a v1 one has no uuid: it is revoked by "all".

What was revoked is kept in each node's store, and every node of a cluster holds all of it: see cluster data sync. A store reset ("reset": true) forgets it on that node, which takes it back from the other nodes when it joins them.

Clients publish and subscribe with topic keys, which a primary client generates with a unitdb/keygen request. The insecure flag of a client's CONNECT, which skips topic key checks, is refused unless the server's config sets "allow_insecure": true, which is for development only and which a cluster node refuses to start with.

A trusted backend, such as an API server acting for its users, needs no topic keys either: give it a service client ID, which only the mintid command issues, with the same key as the server:

> go run ./server/cmd/mintid -config server/unitdb.conf -contract 123456789 -service

mintid reads the keyring as the server does, and mints a v2 ID. Without -contract, it mints a primary client ID of a new contract; -service marks the ID as a trusted service's; -ttl 720h makes the ID expire; -from <id> seals an ID of any key of the keyring, v2 or v1, again as v2 with the issue key, with the same contract, permissions and sessions: it is the only place a v1 ID is still read. mintid mints no v1 IDs since v0.7.0. A service's connections skip topic key checks, in a cluster too. A connection the service opens for a user, with the user's client ID, skips them once the service vouches for it, by publishing {"client_id": "<the service's client ID>"} to unitdb/service on that connection; a connection trusted this way may also generate keys. Keep service IDs on servers, never on clients or devices.

Topics whose first part starts with $ are reserved for the server: no client may publish, subscribe, relay or generate keys for them, a service or an insecure client included. The server keeps its own records under them (see the store's own records).

A session belongs to the client ID that started it: a client of the same contract that sends another client's session key gets a session of its own.

The store's own records

Since v0.7.0 the server keeps its own records under $sys topics, which no client request can address: a contract's subscriptions and the replicas of its messages under $sys.sub.<topic> and $sys.replica.<topic> in the contract's namespace, and the node's own records (hints for other nodes, the topic index, the ids of replicated messages, the security state) under $sys.hint.…, $sys.index.topics, $sys.seen.seen and $sys.security.state in contract 0, which is never a client's. Up to v0.6.0 they were kept under the contract XOR a fixed id, or under a fixed id, so two contracts could share a namespace (one's messages were the other's replicas), and a contract could be drawn equal to a fixed id. Contracts are no longer drawn as 0, the storage engine's master contract, or one of those fixed ids. Sessions and their logs are still keyed by hash, as before.

A store written by v0.6.0 is moved when v0.7.0 opens it, before the node serves: the topic index, the replicas of each topic in it, the ids of replicated messages, the hints and the security state. Each record is copied, the copies flushed, and then the old record deleted, so a crash in between leaves the rest to the next start, which doesn't copy a record twice. Subscriptions aren't moved: they belong to connections, which the restart closes, and clients and the other nodes subscribe again. A wildcard subscription whose first part is * now matches as wildcards do elsewhere in a topic: up to v0.6.0 it matched only a topic whose first part was * itself. See rolling deploys.

Clustering

To bring up the Unitdb cluster start 2 or more nodes. For fault tolerance 3 nodes or more are recommended. Every node needs the same encryption key, or keyring.

> ./bin/unitdb -listen=:6060 -grpc_listen=:6080 -cluster_self=one -db_path=/tmp/unitdb/node1
> ./bin/unitdb -listen=:6061 -grpc_listen=:6081 -cluster_self=two -db_path=/tmp/unitdb/node2

Above example shows each Unitdb node running on the same host, so each node must listen on different ports. This would not be necessary if each node ran on a different host.

Nodes talk over mutual TLS when cluster_config.tls names the cluster's CA and the node's certificate and key: each node needs a certificate signed by the CA, with its node name as a DNS name and for both server and client use, and a tls_addr beside its addr in cluster_config.nodes. A node takes a cluster connection only from a certificate naming another configured node, and refuses a call on it that names another node as its sender. It still listens on its plain addr too, so that a cluster can move to TLS node by node (docs/rolling-deploys.md); set "require": true once every node is on TLS to close it. Until then, firewall the plain cluster ports to the other nodes.

"cluster_config": {
	"nodes": [
		{"name": "one", "addr": "10.0.0.1:12001", "tls_addr": "10.0.0.1:12011"},
		{"name": "two", "addr": "10.0.0.2:12001", "tls_addr": "10.0.0.2:12011"}
	],
	"tls": {"ca_file": "/etc/unitdb/cluster-ca.crt", "cert_file": "/etc/unitdb/one.crt", "key_file": "/etc/unitdb/one.key", "require": true}
}

Client Libraries

Make use of officially supported client libraries to connect to unitdb server running on single node or running on a cluster.

  • unitdb-go Lightweight and high performance unitdb Go client library.
  • unitdb-dart High performance unitdb Flutter/Dart client library.

Architecture Overview

The unitdb engine handles data from the point put request is received through writing data to the physical disk. Data is compressed and encrypted (if encryption is set) then written to a WAL for durability: a Put is written in the background, usually within milliseconds, DB.Flush returns once every entry put before it is written, and a batch returns once its own write is done. Entries are written to memdb and become immediately queryable. The memdb entries are periodically written to log files in the form of blocks.

To efficiently compact and store data, the unitdb engine groups entries sequence by topic key, and then orders those sequences by time and each block keep offset of previous block in reverse time order. Index block offset is calculated from entry sequence in the time-window block. Data is read from data block using index entry information and then it un-compresses the data on read (if encryption flag was set then it un-encrypts the data on read).

Unitdb stores compressed data (live records) in a memdb store. Data records in a memdb are partitioned into (live) time-blocks of configured capacity. New time-blocks are created at ingestion, while old time-blocks are appended to the log files and later sync to the disk store.

When Unitdb receives a put or delete request, it first writes records into tiny-log for recovery. Tiny-logs are added to the log queue to write it to the log file. The tiny-log write is triggered by the time or size of tiny-log incase of backoff due to massive loads.

The tiny-log queue is maintained in memory with a pre-configured size, and during massive loads the memdb backoff process will block the incoming requests from proceeding before the tiny-log queue is cleared by a write operation. After records are appended to the tiny-log, and written to the log files the records are then sync to the disk store using blazing fast block sync technique.

Next steps

In the future, we intend to enhance the Unitdb with the following features:

  • Distributed design: We are working on building out the distributed design of Unitdb, including replication and sharding management to improve its scalability.
  • Developer support and tooling: We are working on building more intuitive tooling, refactoring code structures, and enriching documentation to improve the onboarding experience, enabling developers to quickly integrate Unitdb to their time-series database stack.

Contributing

As Unitdb is under active development and at this time Unitdb is not seeking major changes or new features; however, small bugfixes are encouraged. Unitdb is seeking contibution to improve test coverage and documentation.

Licensing

This project is licensed under Apache-2.0 License.

About

Time-series database and pub/sub messaging server for IoT and real-time devices, in Go: an embeddable storage engine with TTLs, wildcard topics and encryption, and a clustered server over gRPC, TCP and WebSocket with replication and rolling upgrades.

Topics

Resources

Stars

123 stars

Watchers

8 watching

Forks

Releases

Used by

Contributors

Languages