mirror of
https://github.com/featurebasedb/featurebase.git
synced 2026-09-05 08:10:50 +00:00
* unifying idk and featurebase: first pass * resolved conflict with master for gitignore & dockerignore * deleted binaries that were accidentally pushed to git * combined gitlab jobs for idk & featurebase * run go fmt for idk * updated ssh env variable, and made docker password variable in gitlab env variables * fixed typo assigning variable name * trying to fix docker login error * trying a different solution for docker password * pass registry * fixed docker login * updated paths for idk * exclude idk tests from featurebase test run * fix vendor error * update certificates * grpc needs to be in version 1.38 genproto, which is imported by big query updates the grpc version to 1.47.0 grpc 1.47.0 causes etcd to deadlock when calling etcd.Close() the fix is to have a replace in go.mod to specify a specific grpc version * run go mod tidy * go mod * run go mod tidy * exclude bigquery since it is causing issues and undo grpc replace in go.mod * fix grpc version * fix formatting error * update formatting * attempt to fix formatting * update path for code coverage * update to use current branch binaries, not master * fix for building idk - path updates * udpate path for binaries * update job dependecies * update docker idk tests to use the current branch registry * update stages for jobs * updated job dependencies * not allow idk s3 dump to fail since it is a dependency for integration tests * update dependecy for idk tests * update paths for idk build and code coverage * download featurebase binary from s3 * pass branch name to all setup scripts * change to current branch instead of master * updated sonarcloud * sonarcloud fix and branch name fix * trying to speed up pipeline run time * update stage * branch name fix + sonar cloud * sonarcloud |
||
|---|---|---|
| .. | ||
| gen | ||
| testdata | ||
| all-field-types.go | ||
| bank.go | ||
| claim.go | ||
| cmd.go | ||
| common.go | ||
| custom.go | ||
| custom_test.go | ||
| customer.go | ||
| customer_segmentation.go | ||
| customer_segmentation_add_linkedin.go | ||
| customer_segmentation_test.go | ||
| datagen_test.go | ||
| dell.data.go | ||
| dell.go | ||
| dwarranty.go | ||
| equipment.data.go | ||
| equipment.go | ||
| example.go | ||
| hobbies.data.go | ||
| hughes.go | ||
| item.go | ||
| kitchen-sink-keyed.go | ||
| kitchen-sink.go | ||
| locations.data.go | ||
| merck.go | ||
| network.go | ||
| palo_alto.go | ||
| power_scenario.data.go | ||
| power_scenario_1.go | ||
| power_scenario_1_2.go | ||
| README.md | ||
| shared.go | ||
| sites.data.go | ||
| sites.go | ||
| sizing.go | ||
| skills.data.go | ||
| stringpk.go | ||
| texas_health.go | ||
| timeseries.go | ||
| titles.data.go | ||
| transactions.go | ||
| transactions_scenario_1.go | ||
| uscities.data.go | ||
| warranty.go | ||
| zip_codes.data.go | ||
Datagen Tool
Help Usage
Usage of datagen:
-c, --concurrency int Number of concurrent sources and indexing routines to launch. (default 1)
--dry-run Dry run - just flag parsing.
-e, --end-at uint ID at which to stop generating records.
--kafka.batch-size int Number of records to generate before sending them to Kafka all at once. Generally, larger means better throughput and more memory usage. (default 1000)
--kafka.hosts strings Comma separated list of host:port pairs for Kafka. (default [])
--kafka.registry-url string Location of Confluent Schema Registry. Must start with 'https://' if you want to use TLS.
--kafka.subject string Kafka schema subject.
--kafka.topic string Kafka topic to post to.
--pilosa.batch-size int Number of records to read before indexing all of them at once. Generally, larger means better throughput and more memory usage. 1,048,576 might be a good number.
--pilosa.cache-length uint Number of batches of ID mappings to cache. (default 64)
--pilosa.hosts strings Comma separated list of host:port pairs for Pilosa. (default [])
--pilosa.index string Name of Pilosa index.
--seed int Seed to use for any random number generation.
-s, --source string Source generator type. Running datagen with no arguments will list the available source types.
-b, --start-from uint ID at which to start generating records.
-t, --target string Destination for the generated data: [kafka, pilosa]. (default "pilosa")
--track-progress Periodically print status updates on how many records have been sourced.
Example Usage
The following command will create 100 records in Pilosa index (starting at ID 0 and ending at ID 99)
in the equipment index using the equipment data generator.
datagen --source=equipment --pilosa.index=equipment --end-at=99
Adding New Sources
TODO: redo README (or delete?)
If you're looking to add a new Source to datagen, the best thing to do is use the special "custom" datagen source (datagen --source=custom --custom-config=somefile.yaml) and write a somefile.yaml which describes the data you want to generate. An example can be found in datagen/testdata/custom.yaml, and there are some more in the molecula/technical-validation repo.