We centralize the creation paths for test indexes, fields, etcetera so they all have a common path, all using standard test holders. There's still two versions, one for test.* functions and one for internal. They do share a TestHolderConfig though. Large hunks of the related APIs are simplified/streamlined. * Fragments are always created with a Field and don't need a workaround in case they don't have it. * Creation of test fragments, etc., use optional FieldOptions but don't specify names because they're all using new holders for each thing created anyway. This dramatically reduces the complexity of the calls. * test fragments are created inside test views which are created inside test fields, etcetera, so everything is using the same logic; test views aren't bypassing the other layers, they're creating themselves normally within a field. * Quite a few things now use the standard runtime/production logic instead of being custom workarounds; for instance, instead of `mustOpenMutexFragment` creating a fragment and then creating a mutex vector for it, we just create a mutex-typed field and have the normal runtime code do this. * Similarly, we now use the same field creation logic that production does, instead of having our own test-only thing that validates field names directly, so our test that we're validating field names is actually testing the runtime code. Yay. * fragSpec goes away. it was a replacement for fragProxy which existed to solve memory allocation problems but replaced them with interface overhead problems. Now we just have pointers to things and maintain valid data structures. * Many panics are now Fatal or Fatalf calls. * Some specific bugs fixed, like a cluster which was requested and then had its first node directly overwritten, which isn't valid with shared clusters. * Drop the temp-dir test flag and TempDir variable, we can just use $TMPDIR. * Drop a benchmark of "write file to disk" that was purely a benchmark of file write speed, not a benchmark of rendering the data that needs to be written. * Drop the unused "flags" parameter to fragment creation, which was only used back when we changed the BSI format. * Use holder.Txf() rather than index.Txf(). The TxFactory has to be holder-level anyway, referring to it via the index is misleading. * Test holders automatically close themselves and delete themselves, we remove various other things that thought they were responsible for deleting themselves. |
||
|---|---|---|
| .github/workflows | ||
| api/client | ||
| authn | ||
| authz | ||
| boltdb | ||
| client | ||
| cmd | ||
| ctl | ||
| debugstats | ||
| disco | ||
| encoding/proto | ||
| errors | ||
| etcd | ||
| gcnotify | ||
| generator | ||
| gopsutil | ||
| hash | ||
| idk | ||
| ingest | ||
| ingest_testdata | ||
| install | ||
| internal | ||
| lattice | ||
| logger | ||
| lru | ||
| mock | ||
| monitor | ||
| net | ||
| pb | ||
| pql | ||
| prometheus | ||
| proto | ||
| qa | ||
| rbf | ||
| roaring | ||
| scripts | ||
| server | ||
| shardwidth | ||
| short_txkey | ||
| sql | ||
| sql3 | ||
| statik | ||
| stats | ||
| statsd | ||
| storage | ||
| syswrap | ||
| task | ||
| test | ||
| testdata | ||
| testhook | ||
| toml | ||
| tracing | ||
| txkey | ||
| vprint | ||
| .gitignore | ||
| api.go | ||
| api_import_worker.go | ||
| api_test.go | ||
| apimethod_string.go | ||
| audit.go | ||
| audit_internal_test.go | ||
| audit_test.go | ||
| broadcast.go | ||
| bsi.go | ||
| bsi_test.go | ||
| cache.go | ||
| cache_test.go | ||
| catcher.go | ||
| cluster.go | ||
| cluster_internal_test.go | ||
| cmd.go | ||
| CODE_OF_CONDUCT.md | ||
| const_amd64.go | ||
| const_other.go | ||
| dbshard.go | ||
| dbshard_internal_test.go | ||
| dbshard_test.go | ||
| delete_test.go | ||
| diagnostics.go | ||
| diagnostics_internal_test.go | ||
| doc.go | ||
| Dockerfile | ||
| Dockerfile-clustertests | ||
| Dockerfile-clustertests-client | ||
| event.go | ||
| executor.go | ||
| executor_internal_test.go | ||
| executor_test.go | ||
| field.go | ||
| field_internal_test.go | ||
| field_test.go | ||
| filesystem.go | ||
| fragment.go | ||
| fragment_internal_test.go | ||
| gc.go | ||
| gid.go | ||
| go.mod | ||
| go.sum | ||
| hack.go | ||
| handler.go | ||
| holder.go | ||
| holder_internal_test.go | ||
| holder_test.go | ||
| http_handler.go | ||
| http_handler_internal_test.go | ||
| http_handler_test.go | ||
| http_translator.go | ||
| http_translator_test.go | ||
| idalloc.go | ||
| idalloc_test.go | ||
| index.go | ||
| index_internal_test.go | ||
| index_test.go | ||
| ingest_test.go | ||
| internal_client.go | ||
| internal_client_test.go | ||
| iterator.go | ||
| iterator_internal_test.go | ||
| LICENSE-2.0.txt | ||
| license.exceptions | ||
| like.go | ||
| like_test.go | ||
| main_test.go | ||
| Makefile | ||
| metrics.go | ||
| nfpm.yaml | ||
| NOTICE | ||
| pilosa.go | ||
| pilosa_internal_test.go | ||
| pilosa_test.go | ||
| pprof.go | ||
| rbf.go | ||
| README.md | ||
| row.go | ||
| row_test.go | ||
| serializer.go | ||
| server.go | ||
| server_internal_test.go | ||
| server_test.go | ||
| stattx.go | ||
| time.go | ||
| time_internal_test.go | ||
| tracker.go | ||
| tracker_test.go | ||
| transaction.go | ||
| transaction_test.go | ||
| translate.go | ||
| translator_test.go | ||
| tx.go | ||
| tx_internal_test.go | ||
| tx_test.go | ||
| txfactory.go | ||
| txfactory_internal_test.go | ||
| util.go | ||
| util_test.go | ||
| utils_internal_test.go | ||
| version.go | ||
| view.go | ||
| view_internal_test.go | ||
FeatureBase
Pilosa is now FeatureBase
As of September 7, 2022, the Pilosa project is now FeatureBase. The core of the project remains the same: FeatureBase is the first real-time distributed database built entirely on bitmaps. (More information about updated capabilities and improvements below.)
FeatureBase delivers low-latency query results, regardless of throughput or query volumes, on fresh data with extreme efficiency. It works because bitmaps are faster, simpler, and far more I/O efficient than traditional column-oriented data formats. With FeatureBase, you can ingest data from batch data sources (e.g. S3, CSV, Snowflake, BigQuery, etc.) and/or streaming data sources (e.g. Kafka/Confluent, Kinesis, Pulsar).
For more information about FeatureBase, please visit www.featurebase.com.
Getting Started
Build FeatureBase Server from source
- Install go. Ensure that your shell's search path includes the go/bin directory.
- Clone the FeatureBase repository (or download as zip).
- In the featurebase directory, run
make installto compile the FeatureBase server binary. By default, it will be installed in the go/bin directory. - In the idk directory, run
make installto compile the ingester binaries. By default, they will be installed in the go/bin directory. - Run
featurebase server --handler.allowed-origins=http://localhost:3000to run FeatureBase server with default settings (learn more about configuring FeatureBase at the link below). The--handler.allowed-originsparameter allows the standalone web UI to talk to the server; this can be omitted if the web UI is not needed. - Run
curl localhost:10101/statusto verify the server is running and accessible.
Ingest Data and Query
- Run
molecula-consumer-csv \
--index repository \
--header "language__ID_F,project_id__ID_F" \
--id-field project_id \
--batch-size 1000 \
--files example.csv
This will ingest the example.csv file into a FeatureBase table called repository. If the table does not exist, it will be automatically created. Learn more about ingesting into FeatureBase: https://docs.featurebase.com/data-ingestion/enterprise/ingesters
- Query your data.
curl localhost:10101/index/repository/query \
-X POST \
-d 'Row(example=5)'
Learn about supported SQL, native Pilosa Query Language (PQL).
Data Model
Because FeatureBase is built on bitmaps, there is bit of a learning curve to grasp how your data is represented. Data Model Guide: https://docs.featurebase.com/data-modeling-guide/data-modeling
More Information
Installation:https://docs.featurebase.com/setting-up-featurebase/enterprise/installing-featurebase
Configuration: https://docs.featurebase.com/setting-up-featurebase/enterprise/featurebase-configuration
Community
You can email us at community@featurebase.com or learn more about contributing at https://www.featurebase.com/community.
Chat with us: https://discord.gg/FBn2vEp7Na
What's Changed Since the Pilosa Days?
A lot has changed since the days of Pilosa. This list highlights some new capabilites included in FeatureBase. We have also made signficant improvements to the performance, scalability, and stability of the FeatureBase product.
- Query Languages: FeatureBase supports Pilosa Query Language (PQL), as well as SQL
- Stream and Batch Ingest: Combine real-time data streams with batch historical data and act on it within milliseconds.
- Mutable: Perform inserts, updates, and deletes at scale, in real time and on-the-fly. This is key for meeting data compliance requirements, and for reflecting the constantly-changing nature of high-volume data.
- Multi-Valued Set Fields: Store multiple comma-delimited values within a single field while increasing query performance of counts, TopKs, etc.
- Time Quantums: Setting a time quantum on a field creates extra views which allow ranged Row queries down to the time interval specified. For example, if the time quantum is set to YMD, ranged Row queries down to the granularity of a day are supported.
- RBF storage backend: this is a new compressed bitmap format which improves performance in a number of ways: ACID support on a per shard basis, prevents issues with the number of open files, reduces memory allocation and lock contention for reads, provides more consistent garbage collection, and allows backups to run concurrently with writes. However, because of this change, Pilosa backup files cannot be restored into FeatureBase.
License
FeatureBase is licensed under the Apache License, Version 2.0