First, it is possible for us to end up allocating *or freeing* pages during a modification of the free list, in a way such that the change to the free list means that when we finish the modification which caused the allocate or free, we've overwritten the inner change. Second, when deallocating trees, we don't actually deallocate the branch nodes themselves. The former causes potentially severe data corruption. The latter causes us to gradually leak pages in a way that we don't notice because we only run those tests during the RBF tests. The fix for this is surprisingly intricate, because of the counterintuitive fact that *allocating* a page means *removing* things from the free list (and thus potentially deallocating free list pages), while *freeing* a page means *adding* things to the free list (and thus potentially needing to allocate pages for the free list). While modifying the free list, any allocations we need always just come from the end of the file; we don't try to reuse free pages. If a page becomes *deallocated* by a free list modification, we don't annotate it in the free list at the instant that it happens; we stash that information until the current modification of the free list happens, then iterate through any such pages. I am pretty sure there's virtually never more than one, and I don't actually know that I can create a case wherein we'd end up with the nested case firing, wherein removing a page from the free list causes us to remove another page, but I think if the free list got large and cluttered and needed rebalancing or something it could maybe happen. |
||
|---|---|---|
| .. | ||
| cfg | ||
| array.go | ||
| cursor.go | ||
| cursor_internal_test.go | ||
| cursor_test.go | ||
| cursorx.go | ||
| db.go | ||
| db_test.go | ||
| dot.go | ||
| helpers_test.go | ||
| ingest_test.go | ||
| page_map.go | ||
| rbf.go | ||
| rbf_test.go | ||
| README.md | ||
| tx.go | ||
| tx_test.go | ||
| util.go | ||
| util_test.go | ||
Roaring B-tree Format
The RBF format represents a Roaring bitmap whose containers are stored in the leafs of a b-tree. This allows the bitmap to be efficiently queried & updated.
File Format
The RBF file is divided into equal 8KB pages. Each page after the meta page is numbered incrementally from 1 to 2^31.
Pages can be one of the following types:
- Meta page: contains header information.
- Branch page: contains pointers to lower branch & leaf pages.
- Leaf page: contains array and RLE container data.
- Bitmap page: contains bitmap container data.
All integer values are little endian encoded.
Page header
Every page type except the bitmap page contains the following header:
[4] page number
[4] flags (indicates the type of the page)
Meta page
The meta page contains the following header:
[4] magic (\xFFRBF)
[4] flags
[4] page count
[8] wal ID
[4] root records pgno
[4] freelist pgno
Root Records page
A list of all b-tree names & their respective root page numbers are stored in root record pages. Once a bitmap root is created, it is never moved so the root record pages only need to be rewritten when creating, renaming, or deleting a b-tree. If records exceed the size of a page then they are overflowed to additional pages.
[4] page number
[4] flags
[4] overflow pgno
[*] bitmap records
Each bitmap record is represented as:
[4] pgno
[2] name size
[*] name
All bitmap records are loaded into memory when the file is opened.
Branch page
The branch page contains the following header:
[4] page number
[4] flags
[2] cell count
[*] cell index (2 * cell count)
[*] padding for 4-byte alignment
Each cell is formatted as:
[8] highbits
[4] flags
[4] page number
Leaf page
The leaf page contains the following header:
[4] page number
[4] flags
[2] cell count
[*] cell index (2 * cell count)
The leaf page contains a series of cells with the header of:
[8] highbits
[4] flag
[4] child count
[*] array or RLE data
Bitmap page
The data for the bitmap page takes up the entire 8KB.
Proof of Concept Notes
The following are notes made that are temporary for the RBF format. This will change as development progresses:
- Transaction support is deferred
- WAL support is deferred
Pattern of branch splits when adding data in ascending sorted order.
In this example, fan-out is restricted to 2 to make the drawings easy and the splits obvious. The data leaves are only allowed one roaring.Container key (ckey) in this example.
Each frame adds the next datum: A,B,C,D,E,... in order.
The letter represent data leaves, while numbers represent branch pages. The one exception is the first frame where the root is a leaf with data. Every frame after has normal branch at the root.
NB: there are only three places pages get written on addition: a) writeRoot b) putLeafCell c) putBranchCells
add A: root 3 A
add B: root 3 4 5 A B
add C: this sequence of updates occurs
-
putLeafCell writes leaf B to page 5
-
putLeafCell writes leaf C to page 6
-
putBranchCells writes branch page 7 with children 4,5
-
putBranchCells writes branch page 8 with child 6
-
writeRoot writes branch cells to pgno 3, children: 7,8
root 3 7 8 4 5 6 A B C
add D:
root
3
7 8
4 5 6 9
A B C D