We switch everything to use QueryContext/QueryRead/etc instead of Qcx/Tx. We drop the short_txkey subpackage (it's now handled by either keys or querycontext). We drop all the dbshard stuff, and all the tx/txfactory stuff. We remove all the things that related to the old "Block" concept, which was mostly used by the anti-entropy code, but had one fragmentary usage left in the ImportRoaringOverwrite case of ImportRoaring. That's replaced by using a rewriter that deletes all bits (not just bits in specific columns) from an existing thing, but writes in new bits. Actually we could probably do that better with a custom "eradicate-rewriter" that doesn't try to be clever, and just eliminates things. This includes a number of minor bug fixes that were exposed by getting the testing to work. For example: * When checking whether an operation "requires write", we now consider a Delete a kind of a Write, because it is. * Several tests were relying on the fact that writes through Qcx were being committed whether or not the Qcx was ever told to finish. With QueryContext, you actually have to reach a Commit() or the writes don't happen (except for special cases in Delete). * Replaced a lot of panics with t.Fatalf in tests. There's also some minor staticcheck fixes, like deleting the unused "db" member of a boltdb transaction wrapper. |
||
|---|---|---|
| .. | ||
| cfg | ||
| rbf/testdata/check/bad-freelist | ||
| testdata/check | ||
| array.go | ||
| cursor.go | ||
| cursor_internal_test.go | ||
| cursor_test.go | ||
| cursorx.go | ||
| db.go | ||
| db_test.go | ||
| dot.go | ||
| helpers_test.go | ||
| ingest_test.go | ||
| page_map.go | ||
| rbf.go | ||
| rbf_test.go | ||
| README.md | ||
| tx.go | ||
| tx_test.go | ||
| util.go | ||
| util_test.go | ||
Roaring B-tree Format
The RBF format represents a Roaring bitmap whose containers are stored in the leafs of a b-tree. This allows the bitmap to be efficiently queried & updated.
File Format
The RBF file is divided into equal 8KB pages. Each page after the meta page is numbered incrementally from 1 to 2^31.
Pages can be one of the following types:
- Meta page: contains header information.
- Branch page: contains pointers to lower branch & leaf pages.
- Leaf page: contains array and RLE container data.
- Bitmap page: contains bitmap container data.
All integer values are little endian encoded.
Page header
Every page type except the bitmap page contains the following header:
[4] page number
[4] flags (indicates the type of the page)
Meta page
The meta page contains the following header:
[4] magic (\xFFRBF)
[4] flags
[4] page count
[8] wal ID
[4] root records pgno
[4] freelist pgno
Root Records page
A list of all b-tree names & their respective root page numbers are stored in root record pages. Once a bitmap root is created, it is never moved so the root record pages only need to be rewritten when creating, renaming, or deleting a b-tree. If records exceed the size of a page then they are overflowed to additional pages.
[4] page number
[4] flags
[4] overflow pgno
[*] bitmap records
Each bitmap record is represented as:
[4] pgno
[2] name size
[*] name
All bitmap records are loaded into memory when the file is opened.
Branch page
The branch page contains the following header:
[4] page number
[4] flags
[2] cell count
[*] cell index (2 * cell count)
[*] padding for 4-byte alignment
Each cell is formatted as:
[8] highbits
[4] flags
[4] page number
Leaf page
The leaf page contains the following header:
[4] page number
[4] flags
[2] cell count
[*] cell index (2 * cell count)
The leaf page contains a series of cells with the header of:
[8] highbits
[4] flag
[4] child count
[*] array or RLE data or Handle (a pageno) to Bitmap Data
Bitmap Data page
The data for the bitmap data page takes up the entire 8KB.
Proof of Concept Notes
The following are notes made that are temporary for the RBF format. This will change as development progresses:
- Transaction support is deferred
- WAL support is deferred
Pattern of branch splits when adding data in ascending sorted order.
In this example, fan-out is restricted to 2 to make the drawings easy and the splits obvious. The data leaves are only allowed one roaring.Container key (ckey) in this example.
Each frame adds the next datum: A,B,C,D,E,... in order.
The letter represent data leaves, while numbers represent branch pages. The one exception is the first frame where the root is a leaf with data. Every frame after has normal branch at the root.
NB: there are only three places pages get written on addition: a) writeRoot b) putLeafCell c) putBranchCells
add A: root 3 A
add B: root 3 4 5 A B
add C: this sequence of updates occurs
-
putLeafCell writes leaf B to page 5
-
putLeafCell writes leaf C to page 6
-
putBranchCells writes branch page 7 with children 4,5
-
putBranchCells writes branch page 8 with child 6
-
writeRoot writes branch cells to pgno 3, children: 7,8
root 3 7 8 4 5 6 A B C
add D:
root
3
7 8
4 5 6 9
A B C D