mirror of
https://github.com/featurebasedb/featurebase.git
synced 2026-08-28 10:54:59 +00:00
This patch replaces a lot of circumstances in which containers were being copied with circumstances in which they are shared, using copy-on-write semantics. To achieve this, we emulate somewhat the design of go's native `append` function. Operations on a container may optionally yield a new container. A container can be marked "frozen", after which no operation should ever write to it in any way; that applies both to the container itself and the backing store it refers to, if any. So for instance, instead of: c.arrayToBitmap() we now write: c = c.arrayToBitmap() Operations which need to modify a container in any way need to be able to return a new container, which is a modified copy of the previous container. This applies to operations like add/remove, but also to things like unmapping memory-mapped storage, or changing a container's type. Bitmaps do not support the same copy-on-write semantics, currently, but "copying" a bitmap and sharing the containers instead of duplicating them is *much* cheaper than copying the containers. Bitmaps do support a .Freeze method, which currently copies the previous bitmap, making a new one with the same container pointers, and freezes the individual containers. Use this if you need a writeable copy of a bitmap -- the resulting bitmap can safely have its set of containers modified, and bitmap operators that would want to modify the containers will use copy-on-write for that. The primary motivation of this is to reduce the cost of the row cache used by fragments. As a secondary issue, the row cache is no longer updated on writes -- that update was actually a race condition waiting to happen. Rather, writes to a row invalidate the cache entry for that row. The row cache is created by creating a new bitmap, and freezing the relevant containers from the fragment's storage. In the case where nothing is being written, the row cache grows to contain bitmaps containing all those containers, but never copies any containers. If nothing's being read, the row cache is never created, and the containers are in general not getting frozen. The only circumstance where copies have to happen is when things are read (and thus stored in the row cache) and later modified. In that case, each read freezes objects, and the first write to a container after it's been frozen will create a new copy. We drop the enterprise/b btree implementation, because we don't really need it anymore -- we now provide that implementation by default in the open source product anyway. Along with this, there's a lot of other changes which improve support for nil containers, as a cheaper representation for empty containers. Operations which we know will provide an empty container can always short-circuit and just yield a nil *Container. Similarly, operations which would provide a full container can return a single shared full container object (which is frozen). The higher-level (non type-specific) container ops are now using that logic to short-circuit operations for empty and full containers. (For instance, difference of anything minus an empty container is the original thing, union of anything and empty is the original thing, and so on.) The Containers interface adds "Update" and "UpdateEvery" methods, based in part on the "Put" interface provided by the underlying btree implementation; Update performs a possible update in-place of a container for a given key, bypassing the need to replicate the search for that key in the container. UpdateEvery loops through all the containers. Containers do not strictly guarantee that they won't return nil `*Container` objects. However, the container iterators won't return those -- empty containers aren't interesting. Some tests are updated to reflect this. Some of the container internals, like N(), or the isArray() and related functions, accept nil container pointers. Some, like Thaw(), do not. For the array(), bitmap(), and runs() methods, roaringparanoia enables an explicit panic on a nil container explaining the problem, but the intent is that those should never be called unless you already know you have the right kind of container, so by default they don't perform the extra checks. In most cases, this is already covered because a nil container is empty, and there's no operation we can perform that requires us to inspect the contents of an empty container. This is passing a fair amount of testing, but the testing may not be comprehensive enough. The overall impact of this is pretty trivial performance-wise. In our default roaring/ benchmarks, a few things get a few percent faster, or slower. The advantage is that, with read-heavy workloads, the row cache no longer eats up incredible amounts of memory. For a smallish test case, pilosa's memory usage (RES in top) after startup was ~2.5GB. Without this patch, simply reading every row a few times got memory usage to about 9GB, which seemed reasonably stable. With this patch, memory usage went to about 3GB. This will be less noticeable in mixed read/write loads, but it should be consistently significantly lower. In addition to dropping things from the rowCache on modifications, we also stopped performing a full count on a modified row when not using a cache of a kind that would use that count, and don't repopulate the rowCache regardless. We don't want every write to imply a corresponding read after it. There's a lot of room for possible future optimizations in terms of things like in-place operations, and some of the row/rowSegment code is a little suspicious to me, but I don't think it should be *worse* in any cases.
487 lines
12 KiB
Go
487 lines
12 KiB
Go
// Copyright 2017 Pilosa Corp.
|
|
//
|
|
// Licensed under the Apache License, Version 2.0 (the "License");
|
|
// you may not use this file except in compliance with the License.
|
|
// You may obtain a copy of the License at
|
|
//
|
|
// http://www.apache.org/licenses/LICENSE-2.0
|
|
//
|
|
// Unless required by applicable law or agreed to in writing, software
|
|
// distributed under the License is distributed on an "AS IS" BASIS,
|
|
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
// See the License for the specific language governing permissions and
|
|
// limitations under the License.
|
|
|
|
package pilosa
|
|
|
|
import (
|
|
"bytes"
|
|
"fmt"
|
|
"io"
|
|
"sort"
|
|
"sync"
|
|
"time"
|
|
|
|
"github.com/pilosa/pilosa/lru"
|
|
"github.com/pilosa/pilosa/stats"
|
|
)
|
|
|
|
const (
|
|
// thresholdFactor is used to calculate the threshold for new items entering the cache
|
|
thresholdFactor = 1.1
|
|
)
|
|
|
|
// cache represents a cache of counts.
|
|
type cache interface {
|
|
Add(id uint64, n uint64)
|
|
BulkAdd(id uint64, n uint64)
|
|
Get(id uint64) uint64
|
|
Len() int
|
|
|
|
// Returns a list of all IDs.
|
|
IDs() []uint64
|
|
|
|
// Soft ask for the cache to be rebuilt - may not if it has been done recently.
|
|
Invalidate()
|
|
|
|
// Rebuilds the cache.
|
|
Recalculate()
|
|
|
|
// Returns an ordered list of the top ranked bitmaps.
|
|
Top() []bitmapPair
|
|
|
|
// SetStats defines the stats client used in the cache.
|
|
SetStats(s stats.StatsClient)
|
|
}
|
|
|
|
// lruCache represents a least recently used Cache implementation.
|
|
type lruCache struct {
|
|
cache *lru.Cache
|
|
counts map[uint64]uint64
|
|
stats stats.StatsClient
|
|
}
|
|
|
|
// newLRUCache returns a new instance of LRUCache.
|
|
func newLRUCache(maxEntries uint32) *lruCache {
|
|
c := &lruCache{
|
|
cache: lru.New(int(maxEntries)),
|
|
counts: make(map[uint64]uint64),
|
|
stats: stats.NopStatsClient,
|
|
}
|
|
c.cache.OnEvicted = c.onEvicted
|
|
return c
|
|
}
|
|
|
|
// BulkAdd adds a count to the cache unsorted. You should Invalidate after completion.
|
|
func (c *lruCache) BulkAdd(id, n uint64) {
|
|
c.Add(id, n)
|
|
}
|
|
|
|
// Add adds a count to the cache.
|
|
func (c *lruCache) Add(id, n uint64) {
|
|
c.cache.Add(id, n)
|
|
c.counts[id] = n
|
|
}
|
|
|
|
// Get returns a count for a given id.
|
|
func (c *lruCache) Get(id uint64) uint64 {
|
|
n, _ := c.cache.Get(id)
|
|
nn, _ := n.(uint64)
|
|
return nn
|
|
}
|
|
|
|
// Len returns the number of items in the cache.
|
|
func (c *lruCache) Len() int { return c.cache.Len() }
|
|
|
|
// Invalidate is a no-op.
|
|
func (c *lruCache) Invalidate() {}
|
|
|
|
// Recalculate is a no-op.
|
|
func (c *lruCache) Recalculate() {}
|
|
|
|
// IDs returns a list of all IDs in the cache.
|
|
func (c *lruCache) IDs() []uint64 {
|
|
a := make([]uint64, 0, len(c.counts))
|
|
for id := range c.counts {
|
|
a = append(a, id)
|
|
}
|
|
sort.Sort(uint64Slice(a))
|
|
return a
|
|
}
|
|
|
|
// Top returns all counts in the cache.
|
|
func (c *lruCache) Top() []bitmapPair {
|
|
a := make([]bitmapPair, 0, len(c.counts))
|
|
for id, n := range c.counts {
|
|
a = append(a, bitmapPair{
|
|
ID: id,
|
|
Count: n,
|
|
})
|
|
}
|
|
sort.Sort(bitmapPairs(a))
|
|
return a
|
|
}
|
|
|
|
// SetStats defines the stats client used in the cache.
|
|
func (c *lruCache) SetStats(s stats.StatsClient) {
|
|
c.stats = s
|
|
}
|
|
|
|
func (c *lruCache) onEvicted(key lru.Key, _ interface{}) { delete(c.counts, key.(uint64)) }
|
|
|
|
// Ensure LRUCache implements Cache.
|
|
var _ cache = &lruCache{}
|
|
|
|
// rankCache represents a cache with sorted entries.
|
|
type rankCache struct {
|
|
mu sync.Mutex
|
|
entries map[uint64]uint64
|
|
rankings []bitmapPair // cached, ordered list
|
|
|
|
updateN int
|
|
updateTime time.Time
|
|
|
|
// maxEntries is the user defined size of the cache
|
|
maxEntries uint32
|
|
|
|
// thresholdBuffer is used the calculate the lowest cached threshold value
|
|
// This threshold determines what new items are added to the cache
|
|
thresholdBuffer int
|
|
|
|
// thresholdValue is the value of the last item in the cache
|
|
thresholdValue uint64
|
|
|
|
stats stats.StatsClient
|
|
}
|
|
|
|
// NewRankCache returns a new instance of RankCache.
|
|
func NewRankCache(maxEntries uint32) *rankCache {
|
|
return &rankCache{
|
|
maxEntries: maxEntries,
|
|
thresholdBuffer: int(thresholdFactor * float64(maxEntries)),
|
|
entries: make(map[uint64]uint64),
|
|
stats: stats.NopStatsClient,
|
|
}
|
|
}
|
|
|
|
// Add adds a count to the cache.
|
|
func (c *rankCache) Add(id uint64, n uint64) {
|
|
c.mu.Lock()
|
|
defer c.mu.Unlock()
|
|
// Ignore if the column count is below the threshold,
|
|
// unless the count is 0, which is effectively used
|
|
// to clear the cache value.
|
|
if n < c.thresholdValue && n > 0 {
|
|
return
|
|
}
|
|
|
|
c.entries[id] = n
|
|
|
|
c.invalidate()
|
|
}
|
|
|
|
// BulkAdd adds a count to the cache unsorted. You should Invalidate after completion.
|
|
func (c *rankCache) BulkAdd(id uint64, n uint64) {
|
|
c.mu.Lock()
|
|
defer c.mu.Unlock()
|
|
if n < c.thresholdValue {
|
|
return
|
|
}
|
|
|
|
c.entries[id] = n
|
|
}
|
|
|
|
// Get returns a count for a given id.
|
|
func (c *rankCache) Get(id uint64) uint64 {
|
|
c.mu.Lock()
|
|
defer c.mu.Unlock()
|
|
return c.entries[id]
|
|
}
|
|
|
|
// Len returns the number of items in the cache.
|
|
func (c *rankCache) Len() int {
|
|
c.mu.Lock()
|
|
defer c.mu.Unlock()
|
|
return len(c.entries)
|
|
}
|
|
|
|
// IDs returns a list of all IDs in the cache.
|
|
func (c *rankCache) IDs() []uint64 {
|
|
c.mu.Lock()
|
|
defer c.mu.Unlock()
|
|
a := make([]uint64, 0, len(c.entries))
|
|
for id := range c.entries {
|
|
a = append(a, id)
|
|
}
|
|
sort.Sort(uint64Slice(a))
|
|
return a
|
|
}
|
|
|
|
// Invalidate recalculates the entries by rank.
|
|
func (c *rankCache) Invalidate() {
|
|
c.mu.Lock()
|
|
defer c.mu.Unlock()
|
|
c.invalidate()
|
|
}
|
|
|
|
// Recalculate rebuilds the cache.
|
|
func (c *rankCache) Recalculate() {
|
|
c.mu.Lock()
|
|
defer c.mu.Unlock()
|
|
c.stats.Count("cache.recalculate", 1, 1.0)
|
|
c.recalculate()
|
|
}
|
|
|
|
func (c *rankCache) invalidate() {
|
|
// Don't invalidate more than once every X seconds.
|
|
// TODO: consider making this configurable.
|
|
if time.Since(c.updateTime).Seconds() < 10 {
|
|
return
|
|
}
|
|
c.stats.Count("cache.invalidate", 1, 1.0)
|
|
c.recalculate()
|
|
}
|
|
|
|
func (c *rankCache) recalculate() {
|
|
// Convert cache to a sorted list.
|
|
rankings := make([]bitmapPair, 0, len(c.entries))
|
|
for id, cnt := range c.entries {
|
|
rankings = append(rankings, bitmapPair{
|
|
ID: id,
|
|
Count: cnt,
|
|
})
|
|
}
|
|
sort.Sort(bitmapPairs(rankings))
|
|
|
|
// Store the count of the item at the threshold index.
|
|
c.rankings = rankings
|
|
length := len(c.rankings)
|
|
c.stats.Gauge("RankCache", float64(length), 1.0)
|
|
|
|
var removeItems []bitmapPair // cached, ordered list
|
|
if length > int(c.maxEntries) {
|
|
c.thresholdValue = rankings[c.maxEntries].Count
|
|
removeItems = c.rankings[c.maxEntries:]
|
|
c.rankings = c.rankings[0:c.maxEntries]
|
|
} else {
|
|
c.thresholdValue = 1
|
|
}
|
|
|
|
// Reset counters.
|
|
c.updateTime, c.updateN = time.Now(), 0
|
|
|
|
// If size is larger than the threshold then trim it.
|
|
if len(c.entries) > c.thresholdBuffer {
|
|
c.stats.Count("cache.threshold", 1, 1.0)
|
|
for _, pair := range removeItems {
|
|
delete(c.entries, pair.ID)
|
|
}
|
|
}
|
|
}
|
|
|
|
// SetStats defines the stats client used in the cache.
|
|
func (c *rankCache) SetStats(s stats.StatsClient) {
|
|
c.stats = s
|
|
}
|
|
|
|
// Top returns an ordered list of pairs.
|
|
func (c *rankCache) Top() []bitmapPair { return c.rankings }
|
|
|
|
// WriteTo writes the cache to w.
|
|
func (c *rankCache) WriteTo(w io.Writer) (n int64, err error) {
|
|
panic("FIXME: TODO")
|
|
}
|
|
|
|
// ReadFrom read from r into the cache.
|
|
func (c *rankCache) ReadFrom(r io.Reader) (n int64, err error) {
|
|
panic("FIXME: TODO")
|
|
}
|
|
|
|
// Ensure RankCache implements Cache.
|
|
var _ cache = &rankCache{}
|
|
|
|
// bitmapPair represents a id/count pair with an associated identifier.
|
|
type bitmapPair struct {
|
|
ID uint64
|
|
Count uint64
|
|
}
|
|
|
|
// bitmapPairs is a sortable list of BitmapPair objects.
|
|
type bitmapPairs []bitmapPair
|
|
|
|
func (p bitmapPairs) Swap(i, j int) { p[i], p[j] = p[j], p[i] }
|
|
func (p bitmapPairs) Len() int { return len(p) }
|
|
func (p bitmapPairs) Less(i, j int) bool { return p[i].Count > p[j].Count }
|
|
|
|
// Pair holds an id/count pair.
|
|
type Pair struct {
|
|
ID uint64 `json:"id"`
|
|
Key string `json:"key,omitempty"`
|
|
Count uint64 `json:"count"`
|
|
}
|
|
|
|
// Pairs is a sortable slice of Pair objects.
|
|
type Pairs []Pair
|
|
|
|
func (p Pairs) Swap(i, j int) { p[i], p[j] = p[j], p[i] }
|
|
func (p Pairs) Len() int { return len(p) }
|
|
func (p Pairs) Less(i, j int) bool { return p[i].Count > p[j].Count }
|
|
|
|
// pairHeap is a heap implementation over a group of Pairs.
|
|
type pairHeap struct {
|
|
Pairs
|
|
}
|
|
|
|
// Less implemets the Sort interface.
|
|
// reports whether the element with index i should sort before the element with index j.
|
|
func (p pairHeap) Less(i, j int) bool { return p.Pairs[i].Count < p.Pairs[j].Count }
|
|
|
|
// Push appends the element onto the Pair slice.
|
|
func (p *Pairs) Push(x interface{}) {
|
|
// Push and Pop use pointer receivers because they modify the slice's length,
|
|
// not just its contents.
|
|
*p = append(*p, x.(Pair))
|
|
}
|
|
|
|
// Pop removes the minimum element from the Pair slice.
|
|
func (p *Pairs) Pop() interface{} {
|
|
old := *p
|
|
n := len(old)
|
|
x := old[n-1]
|
|
*p = old[0 : n-1]
|
|
return x
|
|
}
|
|
|
|
// Add merges other into p and returns a new slice.
|
|
func (p Pairs) Add(other []Pair) []Pair {
|
|
// Create lookup of key/counts.
|
|
m := make(map[uint64]uint64, len(p))
|
|
for _, pair := range p {
|
|
m[pair.ID] = pair.Count
|
|
}
|
|
|
|
// Add/merge from other.
|
|
for _, pair := range other {
|
|
m[pair.ID] += pair.Count
|
|
}
|
|
|
|
// Convert back to slice.
|
|
a := make([]Pair, 0, len(m))
|
|
for k, v := range m {
|
|
a = append(a, Pair{ID: k, Count: v})
|
|
}
|
|
return a
|
|
}
|
|
|
|
// Keys returns a slice of all keys in p.
|
|
func (p Pairs) Keys() []uint64 {
|
|
a := make([]uint64, len(p))
|
|
for i := range p {
|
|
a[i] = p[i].ID
|
|
}
|
|
return a
|
|
}
|
|
|
|
func (p Pairs) String() string {
|
|
var buf bytes.Buffer
|
|
buf.WriteString("Pairs(")
|
|
for i := range p {
|
|
fmt.Fprintf(&buf, "%d/%d", p[i].ID, p[i].Count)
|
|
if i < len(p)-1 {
|
|
buf.WriteString(", ")
|
|
}
|
|
}
|
|
buf.WriteString(")")
|
|
return buf.String()
|
|
}
|
|
|
|
// uint64Slice represents a sortable slice of uint64 numbers.
|
|
type uint64Slice []uint64
|
|
|
|
func (p uint64Slice) Swap(i, j int) { p[i], p[j] = p[j], p[i] }
|
|
func (p uint64Slice) Len() int { return len(p) }
|
|
func (p uint64Slice) Less(i, j int) bool { return p[i] < p[j] }
|
|
|
|
// merge combines p and other to a unique sorted set of values.
|
|
// p and other must both have unique sets and be sorted.
|
|
func (p uint64Slice) merge(other []uint64) []uint64 {
|
|
ret := make([]uint64, 0, len(p))
|
|
|
|
i, j := 0, 0
|
|
for i < len(p) && j < len(other) {
|
|
a, b := p[i], other[j]
|
|
if a == b {
|
|
ret = append(ret, a)
|
|
i, j = i+1, j+1
|
|
} else if a < b {
|
|
ret = append(ret, a)
|
|
i++
|
|
} else {
|
|
ret = append(ret, b)
|
|
j++
|
|
}
|
|
}
|
|
|
|
if i < len(p) {
|
|
ret = append(ret, p[i:]...)
|
|
} else if j < len(other) {
|
|
ret = append(ret, other[j:]...)
|
|
}
|
|
|
|
return ret
|
|
}
|
|
|
|
// bitmapCache provides an interface for caching full bitmaps.
|
|
type bitmapCache interface {
|
|
Fetch(id uint64) (*Row, bool)
|
|
Add(id uint64, b *Row)
|
|
}
|
|
|
|
// simpleCache implements BitmapCache
|
|
// it is meant to be a short-lived cache for cases where writes are continuing to access
|
|
// the same row within a short time frame (i.e. good for write-heavy loads)
|
|
// A read-heavy use case would cause the cache to get bigger, potentially causing the
|
|
// node to run out of memory.
|
|
type simpleCache struct {
|
|
cache map[uint64]*Row
|
|
}
|
|
|
|
// Fetch retrieves the bitmap at the id in the cache.
|
|
func (s *simpleCache) Fetch(id uint64) (*Row, bool) {
|
|
m, ok := s.cache[id]
|
|
return m, ok
|
|
}
|
|
|
|
// Add adds the bitmap to the cache, keyed on the id. A nil row means
|
|
// deleting the row from the cache.
|
|
func (s *simpleCache) Add(id uint64, b *Row) {
|
|
if b != nil {
|
|
s.cache[id] = b
|
|
} else {
|
|
delete(s.cache, id)
|
|
}
|
|
}
|
|
|
|
// nopCache represents a no-op Cache implementation.
|
|
type nopCache struct {
|
|
stats stats.StatsClient
|
|
}
|
|
|
|
// Ensure NopCache implements Cache.
|
|
var globalNopCache cache = nopCache{
|
|
stats: stats.NopStatsClient,
|
|
}
|
|
|
|
func (c nopCache) Add(uint64, uint64) {}
|
|
func (c nopCache) BulkAdd(uint64, uint64) {}
|
|
func (c nopCache) Get(uint64) uint64 { return 0 }
|
|
func (c nopCache) IDs() []uint64 { return []uint64{} }
|
|
|
|
func (c nopCache) Invalidate() {}
|
|
func (c nopCache) Len() int { return 0 }
|
|
func (c nopCache) Recalculate() {}
|
|
func (c nopCache) SetStats(stats.StatsClient) {}
|
|
|
|
func (c nopCache) Top() []bitmapPair {
|
|
return []bitmapPair{}
|
|
}
|