Indexed columns
Grepr stores each raw log’s attributes as one JSON document and its tags as a map. A search that filters on one attribute reads every log’s whole attribute document to find a match.
An indexed column removes that cost for one field. Grepr writes the field into its own column in the dataset’s data lake table, so a search that filters on the field reads that column and skips the files that can’t contain a match. See The Grepr data lake.
Searches don’t change. The same search returns the same results whether or not a field is indexed, and runs faster when it is.
What Grepr indexes
Each dataset’s raw log table has indexing settings for two groups of fields:
- Attributes, named by their dotted path, such as
http.methodorkube.pod.name. An indexed path has a column for its value, and a path that holds numbers can also have a numeric column, so that range filters skip files too. - Tags, named by their key, such as
regionorpod_name. An indexed tag key has a column for its value.
Grepr always indexes the service and host tags, whatever the settings say.
For each group, you choose whether Grepr discovers new fields, which fields it always indexes, and which fields it never indexes.
Discovery
When discovery is on, Grepr adds a column for a new field it finds in your logs, until the group has as many columns as its discovery cap. Your organization sets the largest cap you can choose. The tag cap doesn’t count service and host.
Discovery is off for a new dataset, so until you turn it on, Grepr indexes only service, host, and the fields on the always-index lists.
When discovery is on and Grepr finds a number at an indexed path, the path also gets a numeric column. A numeric column doesn’t count toward the cap.
Turning discovery off stops Grepr from adding columns. Columns the table already has keep being written. To stop writing a column, add the field to the never-index list.
Always index and never index
Grepr indexes a field on an always-index list whether discovery is on or off, and whether or not the group has reached its cap. Grepr stops writing a field on a never-index list.
For attribute paths:
- A path matches exactly and is case-sensitive. Write the path as it appears in the log.
- An entry that includes nested paths covers the path and every path under it, split at dots. For example,
kubewith nested paths coverskubeandkube.pod.name, but notkubernetes. Wildcards aren’t supported. - An always-index entry marked numeric gets its numeric column the next time a pipeline writes the path, whether or not discovery is on. Use it for values that are numbers written as strings.
- When a path is on both lists, always index wins. To stop indexing a subtree except for one path, add the subtree to the never-index list with nested paths, and add the one path to the always-index list.
For tag keys:
- A key matches exactly.
- A key can’t be on both lists.
service,host, and any key the table is partitioned on can’t be on the never-index list.- To partition a table on a tag key, the key must be on the always-index list, and it stays there while the table is partitioned on it.
When changes take effect
A change reaches running pipelines within about a minute, without a restart. A tag key you add to the always-index list gets its column when you save, so you can partition on it immediately. An attribute path gets its column the next time a pipeline writes a log that has it.
Excluded fields and removed columns
Grepr doesn’t drop an indexed column. When you add a field to a never-index list, Grepr stops writing the field’s column for new logs, and searches read the field from the log’s attributes or tags instead, so they return the same results. The column and the data already written stay in the table, and the column still counts toward the discovery cap. To have columns removed, contact support@grepr.ai.
Change indexing through the API
Reading a table’s configuration returns its indexing settings with a revision. To change them, send the complete settings for both groups to the promotion endpoint, which replaces what’s stored. Include the revision you read: if the settings changed since then, the request fails with a 409 error, and you read them again before you retry. Without a revision, the request replaces the stored settings whatever they are. See the Datasets specification.