Repository navigation
feat(indexes): add algorithm and probe_fraction support - #202
Merged
Merged
Conversation
hotdata-automation
Bot
requested review from
rohan-hotdata
and removed request for
a team
September 29, 2026 03:09
| output_column: Optional[StrictStr] = Field(default=None, description="Custom name for the generated embedding column. Defaults to `{column}_embedding`.") | ||
| vector_precision: Optional[StrictStr] = Field(default=None, description="How precisely a vector index stores each number of a vector. Lower precision shrinks the index so a larger table can be indexed within the same memory, and lets searches run on a smaller instance. Omit this field to store vectors at the same precision as the column, which is the default. The quality figures below come from one benchmark — 1536-dimension text embeddings, cosine distance, default search settings — and are a guide, not a guarantee. Other models, dimensions, distance metrics and data distributions behave differently, so measure on your own data before moving a production index to a lower precision. `float32` — on a `float64` column this halves the index. Widely used embedding models emit 32-bit values, so for those nothing is lost; vectors that genuinely carry more than 32 bits of precision will lose some. `float16` — half the memory of `float32`. In that benchmark its results matched `float32` to within 0.1 percentage points. `float8` — a quarter of the memory of `float32`. In that benchmark it scored about 4 percentage points below `float32`, and raising the search effort did not close the gap, so treat the reduction as permanent for a given index. `float64` — accepted only for a column that already holds double-precision values; it cannot add precision the stored data does not have. Changing this means dropping the index and creating it again. It affects only the index: the table's own values are never altered, and text columns indexed with a generated embedding are not re-embedded.") | ||
| __properties: ClassVar[List[str]] = ["async", "async_after_ms", "columns", "description", "dimensions", "embedding_provider_id", "index_name", "index_type", "metric", "output_column", "vector_precision"] | ||
| probe_fraction: Optional[Union[Annotated[float, Field(le=1, strict=True, ge=0)], Annotated[int, Field(le=1, strict=True, ge=0)]]] = Field(default=None, description="How much of an `ivf` index a search reads, as a fraction greater than 0 and at most 1. Higher finds more of the true nearest neighbours and takes longer. This is a fraction rather than a number of clusters on purpose: the same number of clusters is a different share of the index whenever `nlist` changes, and results would quietly get worse. Omit this for the server's default.") |
There was a problem hiding this comment.
nit: probe_fraction accepts 0 on the client, but the description requires a value greater than 0. (not blocking)
Fix this in the OpenAPI spec (hotdata-dev/www), not in this generated file. Set exclusiveMinimum: 0 on probe_fraction. The generator then emits gt=0 instead of ge=0.
Consequence: probe_fraction=0 passes client validation and the server rejects it. The generated test test/test_create_index_request.py also uses probe_fraction = 0, which is an invalid value by the field description.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Auto-generated from the updated HotData OpenAPI spec.
Source: https://github.com/hotdata-dev/www/pull/436