mirror of
https://github.com/featurebasedb/featurebase.git
synced 2026-08-28 10:54:59 +00:00
31 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e392ce3460
|
enforce int min/max constraints on insert (fb-1772) (#2325)
* moved the debug code to the right spot * enforce int min/max constraints on inserts * add a check for decimal min and max * fixed borked tests * fix the decimal to int conversion in constraint check Co-authored-by: Travis Turner <travis@molecula.com> |
||
|
|
f62313762c
|
implemented extract ddl; tightened up type related stuff (#2329)
* implemented extract ddl; tightened up type related stuff * added some test coverage * review feedback |
||
|
|
2843f218bc
|
Introduce ServiceManager and Refactor DAX Integration tests (#2320)
* Introduce ServiceManager and Refactor DAX Integration tests The ServiceManager provides an interface with which to manage featurebase (dax) services (mds, queryer, computer). It replaces the confusing interface implementations in /dax/server/server.go (which optionally used pointers to in-process objects to satisfy an interface) with (for now) http implementations. The thought is that even if we're running all services in-process, we should communicate between services over http in order to mirror what we would do in a production environment where the services are running on different nodes. This batch of commits does quit a lot, most of which is captured here: - Added `path` support to `dax.Address`. Address is now a string of the form [scheme]://[host]:[port]/[path]. - Added `Holder.directiveApplied` to determine (in tests) if the computer has completed applying the latest directive. This is somewhat temporary until we improve the mds-to-computer logic. - Removed the "service prefix" code which was prepending client URL paths with the prefix. Instead, the serviceType (mds, queryer, computer[n] is now part of `dax.Address`). - Removed, from the dax config, the top level `StorageMethod` and `StorageDSN` and now just have `MDS.Config.DataDir`. - Added `Computer.Config.N` to specify the number of computers to run in-process. - Moved the `pilosa.MDS` interface to `computer.Registrar`. This is an example of getting the interfaces defined in the right packages. - Added `SnapshotTable()` method to the mds client (to align with its API). - Changed `Balancer.AddJob()` to `Balancer.AddJobs()` to support, for example, adding 256 partitions in a single call. Refactored some of the naive Balancer to account for this. - Added a `Seed` to the top-level config. It's not really useful because of package `crypto/rand`. - Added an in-memory implementation of the DisCo interface and disabled etcd in a computer service. - Create sepearte data-dirs for each in-process computer. - Disabled grpc in dax. - Modified the sql3 test definition format to support multiple insert steps and separate query results (to align with those steps). * Changes necessary to get multiple computer instance running in-process For now the config looks like this: ``` [computer] run = true n = 4 ``` but we can probably just change that to be something like: ``` [computer] run = 4 ``` *Issues found running multiple "computers" in-process* - grpc was trying to bind on the same port - changed GRPCListener from `*net.TCPListener` to `net.Listener` - created a nopListener and set to that for now (i.e. disabled grpc) - etcd was starting more than once - changed dax to use in-memory implementations of the disco interfaces (i.e. stop using etcd) - IDAllocator (which uses boltdb) was trying to open the `idalloc.db` file more than once - realized we have to set separate data-dirs for each holder. that fixed it. * Port dax integration tests to ManagedCommand * Modify Balancer-related methods like AddJob to AddJobs There were (and still are) a lot of places where we were adding on job at a time, even when we had a long list of jobs to add. This resulted in every job add (for example adding 1 of 256 shards) taking ~40ms, or over 10s to create a keyed table. One reason was because each job add was making multiple boltdb transactions. * Port over more dax integration test stuff * Add DirectiveApplied to signify that snapshot/writes have loaded. We use this in tests to avoid using sleeps. This should be considered temporary; we're going to need a more robust solution for determining when a computer node is ready to serve complete data. * Finish porting dax integration tests * Improve godocs * Remove docker-based DAX integration tests. * go mod tidy * Move test/managed.go to avoid package conflicts * Modify IDK integration tests to work with ServiceManager changes This is really just computer -> computer0 And the MDS DataDir config change. * cleanup found during review * echo $CI_COMMIT_REF_SLUG in CI * remove docker image arg, use build instead |
||
|
|
0be0c42b66
|
non-sql aggregation, top, decimal and sundries (#2328)
* fixed a bunch of issues with non-pql aggregation; moved some decimal related functionality; made top actually top (for the non-pql case); experimental create function * drive up test coverage |
||
|
|
47d8be26f5
|
implement fb_exec_requests system table (#2327)
implements an fb_exec_requests system table. The purpose of this table is to allow access to internal state to see what queries are running and have been run. Co-authored-by: Travis Turner <travis@molecula.com> |
||
|
|
1ce3a1c103
|
produce a better error when a user tries to sort something unsortable (#2324) | ||
|
|
e0d6b292bc
|
(fb-1779) make optimizer smarter with top operators and aggregate queries (#2323)
* updated optimizer to be smarter when trying to push a top operator down; added test coverage * skip a dax sql test that keeps failing |
||
|
|
3c2c8c6011
|
fixed error messages for alter table add and drop; added test coverage (#2322)
* fixed error messages for alter table add and drop; added test coverage * Removed two CI tests that are failing intermittently for no known reason. |
||
|
|
158cc669d9
|
Tighten up ORDER BY (fb 507) (#2318)
* tighten up checks for order by expressions fixed ordering by expressions * added testing to cover order by cases * Add DecimalAgg member to proto GroupCount definition In DAX, where we have split the orchestrator from the executor, and the orchestrator can run on a different host, there are cases where `GroupCount`s can travel over the wire via the Internal Client. In these cases, when the group count contains a decimal aggregate, we need to send that value as the appropriate type. * fixed missing cases in order by and case block eval Co-authored-by: Travis Turner <travis@molecula.com> |
||
|
|
df829c592f
|
handle filters on _id columns (fb-1765) (#2313)
* handle filters on _id columns using ConstRow * handle keyed and un-keyed _id columns * tests! |
||
|
|
d3f10be743
|
fixed top(x) where top cannot be pushed down into pql query (#2311)
* fixed top(x) where top cannot be pushed down into pql query * review feedback |
||
|
|
f4385df2cf
|
Fix formatting in CLI results with custom SQLResonse.UnmarshalJSON (#2305)
* Fix formatting in CLI results with custom SQLResonse.UnmarshalJSON
When I started this, it was meant to be a quick fix to address the confusing
result formats we were seeing in the CLI. For example, all large integer values
were displayed in scientifc notation. This is because we were passing the result
types from JSON (in this case, float64) into pretty print. Similarly, `IDSets`
and `StringSets` where being printed using the default go Stringer for the types
[]int64 and []string respectively.
I started by writing a customer UnmarshalJSON() method for the `SQLResponse`
type. Part of this (the part which converts data types based on header types)
was already being used in dax tests, so this just formalizes that logic as part
of the `SQLResponse` type.
Then I realized that the sql3 tests (run against the `sql3` package) were
failing because sql3 is not actually returning the `IDSets` and `StringSets`
types. A future task is to formalize return types, define them, and modify sql3
to return them. Once that is done, we can remove the "typed" switch in the
`SQLResponse` json unmarshaller.
Another significant change is the modification to the `ExprDataType` interface:
```
type ExprDataType interface {
exprDataType()
TypeName() string
TypeDescription() string
TypeInfo() map[string]interface{}
}
```
I added two more methods in order to distinguish between a type (`DECIMAL`), its
description (`DECIMAL(2)`), and its type info (`"scale": int64(2)`). Currently,
the description can be used as the field definition in a CREATE TABLE statement,
but we may want to re-think that. Also, Decimal is the only type currently using
TypeInfo.
Finally, I tried to consilidate things around `dax.FieldType` instead of
comparing against parser types outside of sql3. We still have some sql3 parser
and planner types lurking about, but we can address those in future commits.
* Add some test coverage
* smoke test expected INT, now int
* minor fixes
* Introduce WireQueryResponse and related types
This also changes dax.FieldType to dax.BaseType.
* Populate WireQueryResponse correctly
Currently this is in the http handler, and in the queryer.
* Convert sql3 and dax tests to expect pilosa.WireQueryField in results
* fix PQL tests in the SQL defs
* Address a few of the skipped sql tests in dax
|
||
|
|
cca319fbe5
|
Fb 1767 (#2308)
* identifiers can now have the '-' character * skip some integration tests |
||
|
|
5f662d2bce
|
fixed join bugs by fixing query optimizer (fb-1699, fb-1700) (#2307)
this commit changes the way the plan is retrieved; implements Stringer on types.PlanExpression in preparation for HAVING support; removes last vestiges internal float64 arithmetic; implements a filter on PlanOpFilter; fixes various bugs in the PlanOptimizer when rewriting qualified references * fixed selects with unqualified identifiers * handle bad and non-existent query param inputs more appropriately * added test coverage for PlanExpression Stringer Co-authored-by: Matthew Jaffee <jaffee@pilosa.com> |
||
|
|
20a8b5713a |
Add DAX - full list of squashed commits below
In this commit, the Directive is mocked; it doesn't actually reach out to a controller. Limits key translation to only those partitions (per index) specified in the Directive. Attempting to create or find a key (or ID) for a partition which is not handled by this node will result in an error; translation requests are no longer forwarded to other nodes. Limits import into only those shards specified, per index, by the Directive. Attempting to import into a shard which is not handled by this node will result in an error. Stub out /directive endpoint The `applyDirective()` method still needs to be implemented. Update mds references to use the new /mds/types structure In mds, we moved the shared types to mds/types. FeatureBase needs to reference those instead. This also bumps the mds version in go.mod. Implement the Add/Remove Index part of Holder.ApplyDirective() This adds functionality to `Holder.ApplyDirective()` which adds or removes indexes (tables) based on those provided in the Directive. Still to be implemented here: shards and partitions. WIP: remove client from Batch Move Batch into its own package: batch Also, in order to avoid import loops, this introduces packages: /batch/types /client/types Reorganize the Importer-related code Moved the Importer interface to package: batch Move the "pilosa client" implementation of the Importer interface to package: client Modify batch.NewBatch to take an Importer (not client) This commit modifies the batch.NewBatch() function to use a functinal option on Batch to inject an Importer into the Batch. Prior to this, NewBatch() took a pointer to a client, which was a little too restrictive. Now, MDS can implement an Importer which uses information from MDS to determine to which node(s) the import calls should be directed. Add client.SetAuthToken() method to satisfy SchemaManager interface Update ApplyDirective logic to include fields. This needs more work, but it was enough to get a basic test passing. Move Transaction type into /types package. Add interface check on batch.Importer no-op implementation Updated ApplyDirective to create all currently support Field types There are still the following TODOs: - [ ] impolement field options (ex: decimal scale, int min/max, etc). - [ ] `time` fields Added support for Decimal.Scale in ApplyDirective Update mds dependency Add /health endpoint Update to use dax (dax/mds) instead of mds. After moving the mds repository into the dax repository as a sub-package, this commit changes everything in FeatureBase to use the dax repo instead of the now abandoned mds repo. Introduce and implment the WriteLogger interfaces. This adds both a `WriteLogReader` and `WriteLogWriter` interface. They are both implemented by the implementation: `fileWriteLogger`. The `fileWriteLogger` uses the dax/writelogger API to append log messages to files on disk. Add WriteLogWriter.ImportRoaring method to interface This commit adds the `ImportRoaring` method to the `WriteLogWriter` interface. Still to implement are the `Import` and `ImportValue` methods. Reorganize the ApplyDirective code The primary goal was to cache the incoming Directive on the Holder prior to applying all of the changes in the directive (i.e. loading data from the WriteLogger) because applying those changes often validated against the accepted state of the node. Implement all of the WriteLogger read/write methods Implement the HTTP WriteLogger implementation WIP: Introduce shard.Version. Implement snapshotter. Add HTTP Snapshotter implementation This also recofigures server to use the HTTPSnapshotter instead of the FileSnapshotter. Implement snapshotter: TableKeys Implement snapshotter: FieldKeys Dependency dance last of the dependency dance Add support for prototype This adds the Makefile targets to build the docker container and push it to ECR. SQL3 changes which break with dax changes Missed TODO: implement FieldVersion version to WriteLogger Address bug causing missing TranslateStores to error Originally, we tried to limit the TranslateStores which get allocated to only those for which the node is responsible. This works when adding a new table. But if a table already exists, there's no logic to start missing TranslateStores. This reverts back to the old FeatureBase logic which brutishly allocates a TranslateStore for every partition, even if one is not needed. We need to address this by allowing the ApplyDirective logic to initialize TranslateStores when they don't yet exist. Move the ImportRoaringShardRequest type to the types package Since the ImportRoaringShardRequest object is part of the Importer interface, we need to move it to a non-root (i.e. pilosa) package. All the other interface types are either concrete types or part of a sub-package (such as roaring). We do this to prevent an implementer of the interface from having to import the entire pilosa package and risk circular imports. buncha changes to support latest dax stuff Move dax related types to /dax sub-package This commit moves all the common "dax" types into the /dax sub-package. The idea is to ensure that featurebase does not import dax at all. It's ok if dax imports featurebase. In the future, we might need to split the dax sub-package (common data types used by muliple molecula data-plan services) into it's own repo. Add type: dax.Schema This isn't currently being used; I started to use is as a replacement for pilosa_client.Schema, but then deferred that. But we'll need to do it eventually, so it doesn't hurt to have this type in place. Export RowIDs.Merge() method for use in orchestrator. Add CreateSQL method to dax.Table type The CreateSQL() method will return the "CREATE TABLE" statement required to create the dax.Table. Comment out confusing writelogger log message. We need to revisit this, but for now, this log message is confusing. Also, rename daxSharder to versionStore. Remove hard-coded AWS account Implement more FieldOptions such as Epoch Some of the FieldOption logic was stubbed out in the dax package. This commit fills that out more; specifically, it adds the dax.Field.Options.Epoch parameter. export stuff needed for TopK in orchestrator export ValCount stuff to implement Percentile in orchestrator export more stuff to support less code in orchestrator, shared objs Port dax repo over to featurebase/dax (run all as sub-services) This commit does ALOT. Sorry. It introduces a `featurebase dax` sub-command which can be configured to run the various dax services as sub-services within the same process, or individually as the lone service in process. It also changes all the URL paths to be prefixed with the service name. So for example, instead of calling localhost:8080/status, you would now call localhost:8080/featurebase/status. Also, note that all services provide a /health endpoint to confirm they are running in process. Clean up integration tests. Remove PILOSA_ config prefix. Remove duplicate clients (mistake from porting dax to featurebase) Rename sub-service "featurebase" to "computer" In the places where we have hard-coded the sub-service name into a URI path, I've tried to tag the line with a comment containing: `// #SERVICEPATHPREFIX` Update copilot manifest files to reference "computer" Port dax/README.md from dax repository Separate (toml) Queryer Config from Injections We needed to separate the toml config from the configuration required to inject sub-services into the Queryer. I'm not sure this is the best solution, but it's *a* solution. So here we are. Clean up (i.e. remove) the queryer "implementations" package Remove old test file Run WriteLogger and Snapshotter as local sub-services. Prior to this commit, the writelogger and snapshotter services only worked when run as separate services. This allows them to be run in the same process as all the other dax services. There is still some naming issues that we should address, but it's functional for now. Clean up (i.e. organize) the intra-service interfaces. Implement alpha Director for local messages from MDS to Computer Prior to this commit, messages from MDS to the computer service were still going over http. This commit introduces an interface implementation which registers the local computer command, and use that command's API to directly reference methods used by the Director. Clean up a few more interface names Add Queryer OpenAPI document. Update copilot manifests to reflect latest changes Add OpenAPI documents for WriteLogger and Snapshotter Add OpenAPI document for MDS service Add OpenAPI document for Computer service Consolidate errors to use fb/errors package. This commit is a first pass at trying to ensure that all of the DAX code uses: "github.com/molecula/featurebase/v3/errors" This package is a wrapper for "github.com/pkg/errors", so going forward we want to avoid importing that package. The only method which isn't backward-compatible is `New()`; the New() method in the featurebase/errors package takes an errors.Code. If this becomes a problem, we could change this by reverting New() and then introducing something like NewCoded(). But for now I think it might actually discourage someone from just creating a New() error without thinking about how it should be coded. Introduce VersionStore interface Move the existing VersionStore code to the `inmem` package as the in-memory implementation of the new dax.VersionStore interface. Introduce NodeService interface With this, the Controller can maintain a registry of nodes by using this NodeService interface as opposed to an in-memory map of nodes on the Controller struct. This also adds an inmem implementation of the NodeService interface. Introduce controller.Balancer interface This moves the existing balancer package to controller/naive package. The idea is to allow us to add a different Balancer implementation in the future. Introduce DirectiveVersion interface This commit also includes *A LOT* of refactoring to use dax.Worker and dax.Job types everywhere instead of strings. Introduce Schemar interface The previous `Schemar` struct was moved to the `schemar/inmem` package, and `Schemar` is now an interface implemented by that inmem package. Remove unused type `nUnit` Add boltdb implementation of VersionStore interface. This removed the previous sqlite implementation; we decided not to use sqlite for now (as a basic, local disk implementation) because it requires CGO. -------------------------------------------- No longer applicable: Add sqlite implementation of VersionStore interface. This commit implements the VersionStore interface using sqlite. Sqlite requires CGO, so this may not be something we want to include, but it's implemented here to get a feel for how an external implementation might be used; the next step will be to determine how the user configured FeatureBase to run using sqlite as a backing store for services like MDS. Add boltdb implementation of NodeService and DirectiveVersion interfaces. Add boltdb implementation of naive Balancer interfaces. This includes the two interfaces defined in `naive/balancer.go`: - WorkerJobService - FreeJobService Add boltdb implementation of Schemar interface. clean up a linter issue Thread context.Context through all the interfaces. Some of the interface implementations are going to use context, so we need to make that part of the interface. The boltdb implementations, for example, take a context. This is probably so we can do things like cancel or timeout operations. Update interfaces to return error; remove `panic(err)` everywhere. Down-rev grpc version to 1.38.0 Later versions (after 1.42.0?) cause MustRunCluster.Close() in tests to deadlock. This commit also adds an `isComputeNode` feature flag around some of the write log and shard/partition check functionality so that it doesn't run under normal conditions (this is excercised by running the sql3 tests for example). Add MDS_Persistence test to cover meta data persistence This adds a basic test which configures the MDS container to use boltdb as its persistence storage, saved on a docker volume. Then, the mds container is stopped/replaced, and we confirm that the data stored on the volume is availble to the new MDS container. Fix a few things after rebase with sql-experiment branch The lastest version of sql-experiment contains a fairly significan refactor of the way query iteration works. This commit adjusts for those changes. pull dax IDK changes in to FB IDK (#2177) * pull dax IDK changes in to FB IDK * Move docker-related IDK build stuff to featurebase root Building the docker image required the root level go.mod and vendor directory. This change moves the make targets to the root level Makefile, and the Dockerfiles now copy the root level vendor directory (and everything else in the root for that matter). * Fix batch- and client-related tests * InitializePoller on MDS restart/replacement Prior to this change, if MDS was restarted, its internal poller (which maintains an in-memory list of nodes to poll) is empty. This is bad, because it doesn't know about nodes that it should be polling. This change fixes that. Upon MDS startup, it intializes the poller with the list of nodes that MDS keeps in persistent storage (currently: boltdb). * Add EFS volume to MDS Copilot manifest This allows us to use MDS's persistent storage (via boltdb) in the Copilot demo by saving metadata in a boltdb file on EFS. * Thread logger.Logger through all dax components * Revert some of the breaking changes from DAX development. When we first started prototyping DAX, we made changes to the featurebase core code which would have broken the existing featurebase functionality. This commit reverts some of those changes. Anywhere that we need to modify core featurebase functionilty, we put it behind some kind of feature flag. This flag is typically determined by whether the running node is a "compute" node (i.e. DAX.COMPUTER.RUN = true). Co-authored-by: Travis Turner <travis@molecula.com> add packaging for DAX need cgo for datagen build bind to 0.0.0.0, pass GOOS and GOARCH explicitly not sure if the explicit GOOS/GOARCH is actually necessary... Get INSERT INTO (aka ingest) working through SQL3 This commit does a few things which I'll try do describe here. - Introduces a Qctx interface. The existing Qcx is an implementation of this interface, and can be used exactly how it has been. But this allows us to abstract away the notion of Qcx in the Queryer (which is handling SQL3) until we're ready to address that. As an example, the Qcx has a notion of a featurebase Holder, but that doesn't make sense when we're at the Queryer layer. For now, the Qctx used in the Queryer is a no-op. - Adds a ComputeAPI interface implementation for the Queryer. This is effectively the Import() and ImportValues() methods used for ingest. The logic here handles the incoming ImportRequest by first doing any necessary column and row translation for the entire request, then it splits the records by shard, and generates a new ImportRequest per shard with only the shard-appropriate records. - Changes the mds.Importer to take an MDS interface implementation (which can be an mds client) instead of an mdsAddress. This allows us to use a localy MDS implementation rather than assuming we need a client to make calls over a network. Add queryer.Importer interface to handle ingest via SQL (#2203) * Add queryer.Importer interface to handle ingest via SQL This is meant to support ingest through SQL when the queryer and the compute services are running in the same process, or when they are on seperate processes and need to talk via http client. * remove datagen from RPM was originally added as a convenience to generate test data, but is unused and annoying because datagen doesn't easily cross-compile due to cgo * add marshalUnmarshal to controller to avoid passing pointers passing pointers across API boundaries can cause unpredictable things in local vs remote configurations. Co-authored-by: Matthew Jaffee <jaffee@pilosa.com> "fix" a few issues with wrong default partition numbers these still need to be properly fixed and actually get the correct data from MDS go mod tidy Introduce TableQualifier (OrganizationID/DatabaseID) (#2220) * add check in ApplyDirective that version is increasing fix TestAPIDirective to make version always increasing * fix docker image build and break out dax test in CI We have to run the DAX integration tests separately as they call out to Docker, and so it isn't easy to run them in a Docker container as the other tests do. So we run them directly on the CI runner which has Docker and Go installed. We also explicitly exclude these tests from running during the other tests. Also my editor was automatically reformatting some comments badly which is why I added the "data" thing in those two places * add timeout to poller * give Poller a default Logger apparently we can NPE sometimes, seen in CI: https://gitlab.com/molecula/featurebase/-/jobs/3028286364 * bunch of testing fixes, mostly IDK/DAX related make MDS error if sendDirectives errors, don't just log. sendDirectives can error if computer nodes disagree about the validity of a schema (for example), in which case it might need to get deleted and user notified somehow. very messy, needs more thought. re-introduce old env prefix to maintain compatibility with master branch make self-contained dax container for IDK testing build IDK images from source (now that all the source is available since it's in the same repo) catch errors in DoExtractQuery in idktest.go fix IDK bug where prefix path was hardcoded in all cases rather than only when useMDS was true fix TestBatchTargetMDS... needed to add field options and catch error when creating table. also needed an _id field * fix env prefix in tests * WIP getting tests to pass, wanna see CI * don't error if we get a zero version directive and we don't have a directive yet * cleanup debugging junk * "fix" future.rename thing, run IDK tests * Introduce TableQualifier (OrganizationID/DatabaseID) This commit introduces a lot of new types (in dax/table.go) related to TableQualifer (which is made up of OrganizationID and DatabaseID), as well as things like TableID and TableKey. For the most part, we try to thread a QualifiedTableID through the entirety of DAX. There are some places (for example in the Balancers, which are just aware of string keys) which use a string TableKey (tbl__org__db__tableid). * Remove some debugging comments * Add Org/DB support to CLI. This commit adds support for special commands: SET SET ORG acme SET DB db1 USE db1 * remove ".pulled" from IDK Makefile I don't think we need it any more as most things can be built locally. I think it was only there to refresh the FeatureBase images that were tagged as master, but we don't need to do that any more. * Change DAX json tags to kebab-case (i.e. hyphenated) This commit also renames some struct arguments to more accurately reflect their type: for example, renaming `Table` to `TableKey` when the type is TableKey. * Return DAX TableName in SHOW TABLES (instead of Index.Name) There are cases where SchemaAPI is used to return DAX friendly table names (as opposed to featurebase index names, which are DAX TableKey). This is an attempt to do that. With that said, it's not ideal because anything could call those API methods and expect the other type. * Fix a bug which wasn't completely dropping a table. When using boltdb as a backend, DROP TABLE wasn't removing the reverse-lookup key for the table in boltdb. * Remove idk/testenv/certs which got accidentally committed. also update .gitignore to include those. * Fix IDK ingest tests to be TableQualifier aware. * Add example Table types to dax/table.com godoc. * ignore idk.Main fields for flags, upgrade commandeer * go mod tidy * Fix DAX integration tests: ingester using wrong ENV VARs We change from ORGANIZATION_ID to ORG_ID and from DATABASE_ID to DB_ID * Clarify things around idk (docker) tests * Stop running TestKafkaSourceIntegration with t.Parallel() This test can't be run in parallel as it's currently written. Doing so allows for interleaving of messages to the same kafka topic between tests. I didn't attempt to modify the test so it could be run in parallel. That could be done, but left for someone more ambitious. Co-authored-by: Matthew Jaffee <jaffee@pilosa.com> Require Directive.Version be a non-zero value. (#2227) Because the directive cached on the holder is not a pointer, its default version is 0. In order to avoid having to compare against that, we just require that Directive.Version start at 1. General, non-invasive code cleanup and comment adjustment. Move ImportRoaringShardRequest out of the types package Early on in the DAX development, I moved ImportRoaringShardRequest into a types package. There must have been some import loop going on, but since that is not longer the case, it's safe to move this back into the core featurebase (er... pilosa) package. Move Transaction struct back into the pilosa package (from types) Revert some name changes (cli -> client) Add DAX Handler CloseTimeout This was implemented in htt_handler.go, but it had been commented out in the DAX handler. This just uncomments that and finishes the implementation. Remove Qcx from queryer.Importer interface This sets us up to revert the Qctx interface that was initially introduced to allow us to abstract away the need for a Qcx when calling the ComputeAPI from a remote service (i.e. the queryer). Add some go-doc comments and remove unused code. Move SchemaManager setup from datagen to idk.Main (#2233) The set for idk.SchemaManager (for dax implementations) was previously in datagen. This may have been because of some import loop problem during development, but that's no longer an issue. The setup for this should be in idk.Main so anything using that can leverage the MDS-specific SchemaManager setup. Fix issues around nil TxFactory First, don't return a nil. Rather return a new *TxFactory (with no holder). Second, don't call `f.holder` in the testhook outside of checking if `f.holder` is nil. Wrap all bare errors Make service prefixes constants Instead of having `"computer"` throughout the code, use instead a constant: `dax.ServicePrefixComputer`. MDS skip errors when sending empty directives also add in the docker-login and ecr-push changes for serverless DAX Fix the logic in Directive.IsEmpty() (#2236) Update the cached value for Index.translatePartitions In the case where a node already knows about an index, but its assignment of partitions for that index changes (for example, when another node goes down and the node in question is now responsible for more partitions than it previously was), then we need to update the cached value of Index.translatePartitions because that's used in translation checks. minor fixes for IDK-related bugs WIP: tokenize CLI to access cloud FB CLI cloud support with automatic token refresh Also adds support for a GET command which allows making HTTP GET queries to cloud CP which can be handy for debugging stuff. E.g. GET /v2/databases buncha little fixes working on writelogger stuff fix writelogger/snapshotter setup bugs implement writelogging for importRoaringShard add debug endpoint to MDS use shard transactional endpoint in MDS datagen add debugging to API related to writelogger revert handleroption change clean up big PR remove "GET" command from CLI for making arbitrary HTTP request to cloud control plane (was a messy hack and not that useful) remove json tags from FB objects where we had to duplicate the object elsewhere due to import loops and weren't actually json encoding it unexport handlerOption which was exported to try to avoid doing certain things if we're in DAX mode, but I didn't end up merging that code. remove (hopefully) unecessary extra call to api.indexField fix some formatting, unexport some vars, godoc, etc oops, fix build failure Update FeatureBase CLI to support a standard deployment The standard deployment uses a different endpoint and request payload. This commit tries to detect is the standard deployment is being used, and if so, it uses a standard-specific FBQueryer. It also modifies the auto-detection logic to try standard featurebase and dax ports in the case where a port was not provided. MDS API refactor (#2259) * MDS API refactor table IDs are exposed but only created server side also cleaned up dax Makefile * clean up review feedback Co-authored-by: Travis Turner <travis@pilosa.com> * remove TablesByName * rip out inmem implementations and use boltdb everywhere * remove inmem balancer, create bolt tempfile by default on startup * WIP on snapshot table impl and test * Minor comment and code layout adjustments. This commit also adds the `Equals` method to `QualifiedTableID` for equality comparisons. It's no longer safe to compare struct (two structs might still be equal even if one of the structs doesn't have a `Name` value. * Use a unique docker network for each dax test Ocassionally we would see some test failures due to a network already existing. This shouldn't happen, but to avoid that, this commit generates a unique name for each sub test (which gets deleted at the end of every test). * Fix one instance of NewQualifiedTableID losing Name We should probably check the other instances and see if Name is getting lost. * simplify unique network stuff and fix api directive tests * Remove TableIDRequest and TableIDResponse types for /table-id (#2267) For the mds/table-id http requests, just use dax.QualifiedTableID as both the request and response types. * remove lattice from dax, no error on node re-reg, dax docker-compose * various updates * WIP: mds-refactor branch review * no-op on SnapshotTableKeys if table is not keyed * Makefile helpers * add doWeCare so controller doesn't fail unnecessarily * clean up table creation (#2272) * Strip underscores from TableID stub name * fix boltdb versionstore tests: generate unique, sorted tables * fix controller test related to reregistering a node * JobSet -> generic Set Co-authored-by: Travis Turner <travis@pilosa.com> Co-authored-by: Travis Turner <travis@molecula.com> Cleanup after rebase on master The latest rebase on master entailed all the client/batch changes as well as some of the qcx refactoring. It made for a hairy rebase. This commit fixes some of the tests that were failing after that rebase. Fix batch/client import loop missed during rebase (#2280) It's not surprising that `batch` can't import `client`. It was doing that here (importing an error type from the `client` package). What is surprising is that it's okay for `batch_test.go` to import `client` even though `batch_test.go` is an internal test and therefore part of the `batch` package. different boltDB's for schemar/controller, explicit balancers nice helpers for dax docker-compose, make build really fast build FB binary outside of docker, then create Docker image with its working dir in an empty subdirectory so it doesn't send a GB of context to the daemon. error on unassigned jobs and use client with timeout fix CR feedback deregister batch of nodes also make removal faster via director dial timeout implement WorkersForJobPrefix so orchestrator doesn't make up shards also fix some godocs and remove unused method Run sub-tasks of a Directive concurrently in a worker pool. (#2275) * Run sub-tasks of a Directive concurrently in a worker pool. This allows the compute node to concurrently load shapshot and writelog data concurrently, instead of one keyset/partition/shard at a time. It introduces a config parameter called `DirectiveWorkerPoolSize`. * code review cleanup * Use unique container names in DAX integration tests We were seeing "container already exists" errors in CI, so just to be safe, this commit constructs a unique container name for every container in the DAX integration test run. Stub in SystemAPI to Queryer (note: will not work if used) This just makes is so that dax can compile. Actually implementing system-table functionality for dax will take some planning. Tlt/dax merge prep (#2282) * Remove copilot directory * Remove Dockerfile-datagen-long * Remove orphaned RegisterNodeRequest This type is not defined in the dax/mds/http package. * implement TIMEQUANTUM and TTL in Table.Field type * Remove the "service" misdirection in queryer/writelogger/snapshotter. We had originally used an additional layer, er.. package, for a "service". The main distinction was that the Config differed in that it was internal, unlike the Config that we need to provide for the top-level server config (i.e. toml). Having that additional layer just to support a different Config seemed premature at best. So I'm removing it. * Remove dax docker containers no longer used in tests Since we run everything as "featurebase", we don't have multiple container types anymore. * Some minor comment updates * Remove nfpm stuff related to dax * Fix linter issues Fix "duplicate" issues raised by sonarcloud. run docker components of dax integration tests with coverage trying to get dax integration coverage add coverate volume mounts throughout dax integration tests add a lock, tweak dax Makefile, remote flag on query handler remove some unused code convert batch tests to use clustertests to get coverage maybe fix clustertests more authclustertests fixes, test is failing locally but also seems to have been silently failing in CI prior to these changes... let's see if it's still silent fix some lint to kick CI just re-running the job wasn't working... strange behavior remove RetryLogic test and pipe which don't work RetryLogic test removed due to etcd changes. Seebs thinks we shouldn't test this here. Pipe was being ignored since we're no longer using "bash -c" to execute the command. If we need to generate that output file we'll either have to reintroduce bash -c and set -o pipefail so that it actually fails properly, or figure out some other solution. shooting into the dark... first cut at bulk node registration remove unused stuff from batch tests, set coverpkg to ../... batch registration timeout and fix tests disable most tests and don't run fb background batch test debuggin!!!!!!!!! and then he tried this.... Implement importer (for INSERT INTO) in the Queryer Prior to this, we we passing a nil value in for the importer to the planner.NewExecutionPlanner in the Queryer. This meant that INSERT INTO statements didn't work. Now they should. It uses the importer that we build for IDK in /idk/mds/importer.go, and wrapps that with a type that can determine if the provided string "index" is of the form indexName or TableKey. turn off debug mode, fix log saving Run sql3 test definitions in a dax integration test There are currently 22 tests which are not passing. They are skipped in the "skips" slice. WIP, not working, pql queries to tests Add TableQualifier to PQL query logic in the Queryer Add more PQL tests to the keyed table Allow instant node registration if registration-batch-timeout=0 When running dax services in process, we don't want to wait 3s for the compute node to register; we know it's there because it's in the same process. Fixes related to IncludesColumn PQL test. Tests for ConstRow and FieldValue cleanup add UnionRows and Options, better error reporting on bad queries delete unused schemar client.go, clean up unused in batch test CI move test timeouts into more reasonable territory apparently this had already been done, but got merge-stommped at some point move dax bolt test helpers into dax package Add computer CheckIn routine (#2296) * Add computer CheckIn routine This adds a background routine which sends a "check-in" request to MDS every <interval>. This is to address the case where the poller has removed a computer node from the node list (due to a network fault, for example), but the node is still healthy and becomes available again. In that case, the node needs to "check-in" to tell MDS it is still there. MDS will likely send the node a new directive with Method=reset telling the node to delete all of its data an apply the latest directive. * Don't send directives to Deregistered (i.e. removed) nodes We have an issue where we're locking on sendDirective in the controller, and when the node is unavailable, the send hangs and never releases the lock. This is a temporary fix for that until we address the real problem. Fix .gitlab-ci.yml after rebase fix some indentation shenanigans |
||
|
|
8f660a5033
|
SQL BULK INSERT (fb-1749) (#2291)
SQL BULK INSERT This change is to support a BULK INSERT/REPLACE statement that adds the ability to 1) take its input from a file, url or in-line blob 2) to map from the input source to the target columns 3) to transform data (using sql expressions) before inserting 4) support csv and ndjson formats * improving test coverage * increase test coverage again * refactoring for handling transformation with types other than id and int |
||
|
|
eec278893a
|
Move SQL3 test definitions into sql3/test/defs package (#2289)
This so other packages can import and use them in their own tests. |
||
|
|
45067b4e32
|
fb-1744 implement system tables (#2276)
* implement system tables that contain internal state information from FeatureBase * review feedback * review feedback * removed sys prefixes |
||
|
|
0aa5efcc51
|
staticcheck fixes (#2278) | ||
|
|
193ef7cba1
|
Test expression eval for tuple values in inserts (fb-1555) (#2270)
refactored comparison, equality and arithmetic expr eval for decimal data types and added a test to cover expression eval for inserts fixed failing test |
||
|
|
00ef2380e5
|
Batch insert via SQL (multiple tuples) (#2243)
* Formatting adjustments made during code review. While reviewing the BULK INSERT logic (in order to decide how best to approach "ingest via sql" in the cloud), I made a few formatting and comment changes. I'm just adding them here as a separate commit so they don't muddy up my actual work. * Parser modifications to support mulitple tuples in INSERT INTO This commit doesn't include all of the changes required in the planner. Fow now, the planner is simply modified to continue supporting a single tuple (the first tuple in the list). * Update the planner to handle multiple INSERT INTO tuples This is part 1. It's still using the existing logic which builds an ImportRequest for every record (and every field!). The next step will involve using a client.Batch to handle the records. * Introduce client.Importer interface (used by client.Batch) Instead of the Batch having a pointer to a client, this puts an interface there instead (which the client implements). It also allows us to inject a different importer (i.e. other than a featurebase.client) into the Batch. * Decouple batch from client This commit pulls batch-specific code out of the client package and into a new batch package. It introduces the batch.Importer interface, the methods of which replace all the calls that batch was previously making directly to client methods. Finally, it contains two implementations of the batch.Importer interface: one is a wrapper around client, and the other is a wrapper around featurebase.API. * Use docker (instead of MustRunCluster) for internal batch tests Because the `batch` package tests are internal, using test.MustRunCluster() resulted in an import loop (because it eventually imports `server`, and we can't have that). So this commit replaces the use of `test.MustRunCluster()` with docker. The setup is basically the same as that used in the idk docker tests. Here we also remove all client-side references to `UseIngestAPI`, which is an experimental (json) ingest api. It's still suppored on the server, but here we remove the external usage of it. * cherry-pick fix * Use batch.Import() for sql3 INSERT INTO statements * Thread logger into sql3 * fix batch test * Fix some shadowing complaint by linter * Address some test issues related to stringsets * Exclude batch integration tests from CI * Address PR feedback - Added description to batch.README - Consolidated grep commands in .gitlab-ci.yml - Removed some debugging comments - Replaces some inadvertantly removed license headers * Add batch package to gitlab CI * Updated CI for batch package Updated CI include path Update gitlab ci Update CI Update CI Trying new include path for ci Updated gitlab ci include path Made idk race job optional for sonarcloud upload add testdata directory remove testenv from dockercompose file use GIT_STRATEGY clone in batch CI add testdata volume to dockercompose Co-authored-by: Fletcher Haynes <fletcher.haynes@generalassemb.ly> |
||
|
|
e1b704d4ea
|
added a test to cover the keyword replace as being synonymous with insert (#2261) | ||
|
|
b32c33c992
|
tightened up is/is not null filter expressions (FB-1741) (#2260)
Covers tightening up handling filter expressions that contain is/is not null ops. These filters may have to be translated into PQL calls to be passed to the executor and even though sql3 language supports nullability for any data type, currently only BSI fields are nullable at the storage engine level (there is a ticket to add support for non-BSI field here FB-1689: IS SQL Argument returns incorrect error) so when these fields are used in filter conditions we need to handle BSI and non-BSI fields differently. |
||
|
|
fb40cdc2cb
|
fb-1729 Enriched Table Metadata (#2255)
enriched metadata for tables added support for the concept of a table and field owners in metadata; mechanism to derive owner from http request metadata; metadata for table description |
||
|
|
a8146f826d
|
moved file from this repo to documentation repo (#2232) | ||
|
|
1a6c15263d
|
And now....INNER JOIN! (#2230)
* ID sql3 internal type representation is int64; fixed a bug that assumed incorrectly that it wasn't * refactored some names for clarity * primary: get nested loop joins to work; secondary get brute force aggregations for SUM working * added tests; removed debug output * review feedback * Update sql3/planner/compileselect.go review feedback Co-authored-by: Travis Turner <travis@pilosa.com> Co-authored-by: Travis Turner <travis@pilosa.com> |
||
|
|
9c216ef8a6 |
use testhook test cleanup
The testhook post-test hooks only work if you use a TestMain to invoke them, otherwise the cleanups can be registered but never actually get run. This deletes the etcd sockets, and temp directories, that we created from our test runs. We also fix the test creating a temp file directly to create it in a TempDir (which gets cleaned up after the test), and fix the name of the top-level tests displayed in TestMain. |
||
|
|
91e3b8457a
|
fb-1075 (#2221)
* handle multi field count correctly COUNT() should ignore null values. If the data type of the expression supports an existence bitmap for the underlying FeatureBase data type we will use it to eliminate nulls from the aggregate * simplify aggregate for existence test we can use a direct != null instead of an indirect not(=null), and avoid relying on the probably-broken behavior in the executor that tries to silently fix up Row(x=3) tests on BSI fields which wanted Row(x==3). Co-authored-by: Seebs <seebs@molecula.com> |
||
|
|
7c936471cc
|
sql3 changes (#2211)
* first cut of working (slowly) bulk insert; table valued functions and a tuple data type to support time quantums * oversight * filter pushdown implementation; bulk insert * addressed some linter issues |
||
|
|
2052bb01d8 |
refactor testing to share clusters more often
When doing tests, we create a ton of one-off clusters. This turns out to be expensive and slow. Fixing it is surprisingly hard. Fundamentally: If we're sharing clusters, we need to use different indexes for each test, to avoid clashes. This changes index names. As a side-effect, this reorders many partition-based things, like the order keys are returned in. Thus, to fix this, we change a lot of tests to no longer depend on the *order* in which strings are returned. Having done that, we can also discard the ModHasher behavior, since that only existed to allow us to reliably predict partitioning. The basic design is as follows: Instead of a cluster being a []*Command, a "shareable" cluster is now a []*Command plus some flags, and a "cluster" is a pointer to a possibly-shared cluster, plus a link to the specific test using this specific cluster, and correspondingly, its test name suitably coerced to be a valid index name prefix. The "test.Cluster" object now has methods to allow retrieving an index name, and also implemnts fmt.Formatter to let you use, e.g., `%i` with it in Sprintf to get "the index name, plus an i". (This works for everything but %p and %T.) This allows us to consistently rework all the many things that use index names in a persistent way. We also have `MustUnshared` and `MustRunUnsharedCluster` methods which allow us to specify that a given test needs its own cluster for some reason. For instance, the tests that want to run backups need their own isolated cluster, and the tests that want to close or reopen nodes need their own cluster because a reopened cluster won't have working GRPC for some reason. On "closing" a shared cluster (actually the test-specific wrapper that reflects a given sharing), we delete any indexes starting with that test's index name prefix. Otherwise, the huge pile of open indexes prevents `go test -race` from working on MacOS, where we run out of address space too quickly. This is fairly enormous but most of the individual changes are fairly trivial things like replacing the string "i" with "c.Idx()". We also tweaked a test that failed for me a couple of times to not depend on sort order. |
||
|
|
e4a4a06af0
|
added /sql endpoint; implemented SHOW TABLES (#1935)
* squashed 45 commits into one :) * tlt/sql experiment (#2035) * Move PlanOperator to sql3/planner/types package includes: type PlanOperatorColumn struct type PlanOperator interface * Remove planner dependencies from pilosa package The goal after this is to prevent the planner package (which doesn't exist yet) from being imported by the pilosa package; we just want it injected into the server in server/server.go. This is because the planner package uses pilosa types, so we need to avoid circular dependencies. Added ExecutionPlannerFn Make public: pilosa.ExecOptions Added a pilosa.Executor interface Added a planner.types.CompilePlanner interface Isolated the planner calls to: - Executor.Execute() - *API.[method]() * Move executionplanner files into the sql3/planner package. This required a bit of gymnastics, and there are some things around FieldOptions which need to be addressed soon. * Remove the hacky FieldOptions stuff I added earlier This implementation just uses the pilosa.FieldOption functional options provided by the API (as opposed to trying to build a FieldOptions object. It also changes field types to constants. These are private for now, but if we need to make them public, we should put them in the planner/types package. * Implement the "scale" value from Decimal(scale) Also, precision and scale were currently reversed in the parser. This fixes that. * Modify the parser to handle CACHETYPE <type> SIZE <size> It's a little odd to me that the cache type values are Tokens, but I guess it's ok. One thing to keep in mind is that FeatureBase expects lowercase values, so this commit changes the parser to set the value to the lowercase version of the type. * Fix the /sql2 tests This entailed a combination of commenting out or t.Skip()-ing tests which covered code in the parser that has been commented out or removed as not currently supported in sql3. It also adds some coverage for the sql.Contraint stringers. * Prevent JSON sql results from containing closing commas This commit just re-works the existing output code to avoid inserting closing commas (which results in invalid JSON). * Enhance the CREATE TABLE test coverage. In particular, ensure that the fields which get created in FeatureBase are what we expect based on the fields defined in the CREATE TABLE statement. This also ensures that the TIMEQUANTUM and CACHETYPE contraints are not provided for the same field (since those constraints are not supported together). * Adjust the EBNF file to indicate SIZE contraint is optional A CACHETYPE can be provided without a SIZE. This change indicates that SIZE is optional. * Remove `executionplanner_` from file names (#2040) * implementation of ALTER TABLE (sans column RENAME) * refactored expression analysis; added more robust type checking; all unary and bin ops function on ints * added type support for expressions; full bin/unary op support; added cast; more literal support * cast int to all other types * all literals (except idset, stringset & timestamp) make it thru; cast to all types with int as source now works * implemented LIKE/NOT LIKE * Implemented IS [NOT] NULL * Move sql2 files into sql3/parser package (#2045) * Move sql2 files into sql3/parser package This also removes the sql2 package. * Fix tests which were typing _id fields as INT intead of ID * implemented BETWEEN, NOT BETWEEN * Add featurebase/error package (#2046) * Add featurebase/error package I copied the `dax/errors` package which I am starting to use in the DAX prototype into `featurebase/errors` in order to start using it with the sql3 package. It's basically a wrapper around `github.com/pkg/errors`, but it uses a customer coded error. The sql package can define its own errors based on the `featurebase/errors` types. Then do things like `Wrap()` and `Is()`. * Address the linter complaints: shadowed variables, unreachable code * implemented IN & NOT IN with expression lists * first cut of CASE * Fixed some errors from rebase * updated bnf; removed unused code; tightened up error handling * first crack at basic CLI for SQL3 Use: `featurebase cli` Still lots to do here, but for example: > select count(*) from tremor +--------------+ | COUNT | +--------------+ | 1.158321e+06 | +--------------+ * Iterate on the CLI (#2057) Handle the errors. Add an "exit" command. Add some general formatting and white space. Add termination character: ";" (semicolon) This commit allows a user to provide multiple or partial SQL statements. Example of multiple statements: ``` show tables; select * from foo; ``` Example of partial (multi-line) statements: ``` select * from foo; ``` Don't uppercase the header values * error refactoring; first cut of TOP; remove unused code; use log.Printf instead of fmt.Printf * fixed a bug with QualifiedRef from refactoring; added bones of INSERT; removal of unused code; tightened up errors more; fixed failing tests * single value list for INSERT * Update bnf per discussion with Travis; INSERT now doing the requisite stuff * Pat's eyes went square - nothing wrong with TOP, Pat needed to learn arrays again. * improved some errors; fixed tests to suit * send warnings back in the api; update CLI to display warnings * start warning on stuff not implemented so we don't get bugged about it * Tlt/sql experiment (#2063) * Expresssion -> Expression * Add SQL planner test - adds a test to which it is easier to add tables and SQL statments - un-exports all of the expression types - removes the planner pointer from the expression types (it can be added back later if need be) * Fix where clause on a string field Prior to this commit, the binary expression for a where clause on a string field was building the call by providing a range operator which is typically used for BSI fields. This changes it to use the call.Args for string values. * Update planner tests to handle multiple sql for the same results * Reorganize SQL tests Introduce a test/helpers package and move shared MustQueryRows into that package. * Add a compatibility map for field types. (#2064) This is primarily to address the fact that ID fields were previously incompatible with INT literals. We should probably consider introducing a custom type for FieldType which can be used to define compatibilities. * significantly refactored type checking * Handle nil (NULL) values in the sql CLI. (#2067) go-pretty panics if the interface{} field value is nil. This replaces nil values with a "NULL" string. * Squash some commits fixed a still failing test added line, col to all error messages refactored source handling to enable table aliases fixed some copypasta per review warnings for order by & topn; implemented select as a source starting to handle in (select...); added stub for optimizer JSON-encode the sql error and warning strings (#2069) Error strings with unencoded characters (like double quotes) were resulting in invalid json. got insert working; added symbol table; added concrete optimizer; added nascent NestedLoopsOperator; rewrite "where foo in (select..." as inner join * all about the sets (#2085) * implemented setcontains() * implemented set literal; insert set column values; setcontains/all/any both in expr eval and pql filters * Convert test to use latest framework. (#2086) * fixed some comments * removed refactored tests Co-authored-by: Travis Turner <travis@pilosa.com> * Add support for Decimal fields to the sql test. (#2090) * dates (#2094) * return dates as strings in output; tightened up decimal type checking * return dates as strings in output; tightened up decimal type checking * fixed failing tests after decimal changes * can now insert decimal values * implemented insert for timestamp data type; implemented current_date, current_timestamp constants * fixed some failing tests * handle date literals from strings in insert statements * changes from feedback * Fix pointer method error * sql3 API interface (#2110) * Introduce API-related interfaces: SchemaAPI, ComputeAPI The sql3 code was relying on the pointer: *pilosa.API in order to call API methods directly on the local node. If we want to import and use the sql3 package in another service (the DAX queryer, for example), we need to be able to use an implementation of an interface for those API method calls. This commit introduces two interfaces, both automatically implemented by pilosa.API: - SchemaAPI - ComputeAPI * Convert sql3 code to use IndexInfo instead of Index The sql3 code was relying on a *pilosa.Index and its methods to get general information like index and field name, type, etc. This commit converts everything to use a *pilosa.IndexInfo instead. This allows us to modify the SchemaAPI interface to also return IndexInfo instead of Index, which will be a lot easier to implement in a non-pilosa package (like DAX); creating a *pilosa.Index requires providing things like data directory paths and holders, which are not necessary for these use cases. * Unary and Binary Ops R US plus CAST (#2111) * implemented string literal for timestamp epoch * fixed failing test * fixed the failing test again * refactored tests; implemented unary op tests for all datatypes; implemented binop tests for int/int, int/id, int/decimal & ID/int * implemented all binary ops for INT & all other types, ID & all other types * implemented binary ops for DECIMAL types & all other types * added STRING & BOOL to various tests; implemented all remaining binOp tests * fix up some stuff after rebasing * refactored test defs into multiple files; implemented CAST for every datatype * added tests for like/not like * addressed review feedback * addressed type review feedback * tightened up IS [NOT] NULL behavior plus tests (#2118) * tightened up IS [NOT] NULL behavior plus tests * BETWEEN/NOT BETWEEN with all data types * addressed review feedback * Handle negative integers in column min/max constraints (#2120) This commit parses the min/max contraint as an expression, as opposed to an int literal, so that negative values are treated as Unary expressions. There currently isn't support for min/max constraints on `decimal` fiels, so for now this change only expects +/- integer values. * Implement the CREATE TABLE keypartitions logic (#2123) * Execution time, IN/NOT IN & multiple aggregates (#2124) * added display of execution time * IN/NOT IN tests for all data types * fixed date parsing * removed duplicative tests * refactoring aggregates * suport multiple aggregates * Address review feedback * final round of feedback * Add method SchemaAPI.CreateIndexAndFields() (#2127) In order to support a CREATE TABLE statement as a single command, this commit alters the SchemaAPI interface to contain a single method which handles both the index and its fields. It also updates the sql3 code to use this interface instead of CreateIndex() and CreateField() indepedently. * Symbol Handling (Again) (#2129) * Refactored symbol handling in the planner; re-instated the select as source tests * removed commented out code * addressing review feedback * Move hard-coded _id field out of planner and into interface implementation (#2130) This commit moves the hard-coded addition of the `_id` field from the planner to the SchemaAPI.IndexInfo() implementation method. NOTE: If anything was expecting SchemaAPI.Schema() to also return the `_id` field as part of its field list in each table, then it would not be there because the `_id` field is only added in the IndexInfo() method for now. Currently that's not a problem because nothing is expecting the `_id` field for `Schema()`. * Multiple aggregates, all aggregates stand alone and in GROUP BY (#2132) * handle multiple aggregates in group by queries * added handling for avg() aggregate both stand alone and in group by * tightened up sum & avg outside of group by * added min, max & percentile * added warnings * Make MaterializedRowSet implement the PlanOperator interface. (#2133) This commit refactors the PQLMultiGroupByOperator to have a PlanOperator as its output. Then, when it initializes, it sets up a MaterializedRowSet and populates that with the values from the multiple group by operations. * added explicit min/max pql operators * saved a file I forgot to save * per review * Un-indent some if/else nesting (#2136) Co-authored-by: Travis Turner <travis@pilosa.com> * Add optional `name` argument to test structs. This commit adds the `name` argument to `tableTest` and `sqlTest` so that a test can be optionally named. This allows a developer to more easily run/identify a particular test by name. * Inbuilt functions (redux) (#2141) * set functions type parameter type checking * implemented datepart * include SQL3 type in SHOW COLUMNS output * fixed select as source; failing SHOW COLUMNS test * select in select list * dump output columns; handle optimization for select list subqueries * make it an error to return multiple rows for a select list subquery * added description * contants and test coverage for datepart function * SQL3 Refactor-palooza (#2182) * removed unneeded IsAggregate() * first cut of working nested loops operator aka INNER JOIN * remove selectListItemPlanExpression * added some warnings * all the tests are passing again! * addressed some linter complaints * added basic order by * bug fixes; added 'or replace'/'replace' to insert * for insert references should return appropriately * added back ability to use subquery singleton expressions * removed dead code; fixed test * json-able plan, Schema() plus refactoring * fixed dumb code * add some tests for time quantum behavior * Code cleanup during review. Also fixed INSERT to keyed table bug. This commit contains a lot of minor adjustments made during code review. It also contains a bug fix that was preventing INSERT into a keyed table (i.e. _id type STRING) from working. Co-authored-by: Travis Turner <travis@molecula.com> * Fix expected min/max on timestamp column test (decimal field) I don't know why this changed, but presumably something to do with decimal related work that happened on master. * Fix compile problem after rebase * review feedback Co-authored-by: Matthew Jaffee <jaffee@pilosa.com> Co-authored-by: Travis Turner <travis@pilosa.com> Co-authored-by: Travis Turner <travis@molecula.com> Co-authored-by: Fletcher Haynes <fletcher@capitalprawn.com> |