featurebase/idk/datagen
Matthew Jaffee 3693b9950a
Expose translation mvcc (#2381)
* expose Transaction on TranslateStore for DAX Snapshotting

* try to fix ramdisk nonsense

apparently, we were running in either a shell env or docker env
randomly, so this could sometimes pass and sometimes fail since the
shell env had the ramdisk set up and docker didn't.

Now we force to run in docker always and set up ramdisk explicitly.

* ramdisk mount should be defined on gitlab runner config now

* debug ramdisk issue?

* fix tests... and a buncha other stuff

Took retry out of CI config because I think it's doing more harm than
good at this point.

The executor test I modified failed when I changed DefaultPartitionN
to 8, but just because stuff was out of order so I made it more
robust.

I edited some data gen stuff to make shorter lines because it was
making grep results unusable.

the actual fix is in translate_boltdb_test.go

* clean up, fix code review feedback
2022-12-20 07:45:32 -06:00
..
gen FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
testdata Add DAX - full list of squashed commits below 2022-11-16 11:19:40 -06:00
all-field-types.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
bank.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
claim.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
cmd.go Add DAX - full list of squashed commits below 2022-11-16 11:19:40 -06:00
common.go staticcheck fixes (#2278) 2022-11-07 10:51:55 -06:00
custom.go Add DAX - full list of squashed commits below 2022-11-16 11:19:40 -06:00
custom_test.go staticcheck fixes (#2278) 2022-11-07 10:51:55 -06:00
customer.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
customer_segmentation.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
customer_segmentation_add_linkedin.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
customer_segmentation_test.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
datagen_test.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
dell.data.go Expose translation mvcc (#2381) 2022-12-20 07:45:32 -06:00
dell.go staticcheck fixes (#2278) 2022-11-07 10:51:55 -06:00
dwarranty.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
equipment.data.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
equipment.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
example.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
hobbies.data.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
hughes.go staticcheck fixes (#2278) 2022-11-07 10:51:55 -06:00
item.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
kitchen-sink-keyed.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
kitchen-sink.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
locations.data.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
merck.go staticcheck fixes (#2278) 2022-11-07 10:51:55 -06:00
network.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
palo_alto.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
power_scenario.data.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
power_scenario_1.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
power_scenario_1_2.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
README.md FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
shared.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
sites.data.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
sites.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
sizing.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
skills.data.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
stringpk.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
texas_health.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
timeseries.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
titles.data.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
transactions.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
transactions_scenario_1.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
uscities.data.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
warranty.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00
zip_codes.data.go FB-1597: unifying idk and featurebase (#2160) 2022-07-28 17:23:16 -05:00

Datagen Tool

Help Usage

Usage of datagen:
  -c, --concurrency int             Number of concurrent sources and indexing routines to launch. (default 1)
      --dry-run                     Dry run - just flag parsing.
  -e, --end-at uint                 ID at which to stop generating records.
      --kafka.batch-size int        Number of records to generate before sending them to Kafka all at once. Generally, larger means better throughput and more memory usage. (default 1000)
      --kafka.hosts strings         Comma separated list of host:port pairs for Kafka. (default [])
      --kafka.registry-url string   Location of Confluent Schema Registry. Must start with 'https://' if you want to use TLS.
      --kafka.subject string        Kafka schema subject.
      --kafka.topic string          Kafka topic to post to.
      --pilosa.batch-size int       Number of records to read before indexing all of them at once. Generally, larger means better throughput and more memory usage. 1,048,576 might be a good number.
      --pilosa.cache-length uint    Number of batches of ID mappings to cache. (default 64)
      --pilosa.hosts strings        Comma separated list of host:port pairs for Pilosa. (default [])
      --pilosa.index string         Name of Pilosa index.
      --seed int                    Seed to use for any random number generation.
  -s, --source string               Source generator type. Running datagen with no arguments will list the available source types.
  -b, --start-from uint             ID at which to start generating records.
  -t, --target string               Destination for the generated data: [kafka, pilosa]. (default "pilosa")
      --track-progress              Periodically print status updates on how many records have been sourced.

Example Usage

The following command will create 100 records in Pilosa index (starting at ID 0 and ending at ID 99) in the equipment index using the equipment data generator.

datagen --source=equipment --pilosa.index=equipment --end-at=99

Adding New Sources

TODO: redo README (or delete?)

If you're looking to add a new Source to datagen, the best thing to do is use the special "custom" datagen source (datagen --source=custom --custom-config=somefile.yaml) and write a somefile.yaml which describes the data you want to generate. An example can be found in datagen/testdata/custom.yaml, and there are some more in the molecula/technical-validation repo.