Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -34,7 +34,7 @@ DATA = sql/$(EXTENSION)--$(EXTVERSION).sql \
sql/$(EXTENSION)--1.0-beta3--1.0.sql

# Test configuration for pg_regress
REGRESS = setup chunking multibyte_chunking hybrid_chunking queue delete_truncate delete_truncate_pk pk_type_session max_retries vectorization multi_column maintenance edge_cases providers worker cleanup embedding pk_types stale_embeddings hybrid_test count_tokens vectorizer_status per_table_model
REGRESS = setup chunking multibyte_chunking hybrid_chunking queue delete_truncate delete_truncate_pk pk_type_session max_retries vectorization multi_column maintenance edge_cases providers worker cleanup embedding pk_types stale_embeddings hybrid_test count_tokens vectorizer_status per_table_model model_drift
REGRESS_OPTS = --inputdir=test --outputdir=test

# Documentation files (if any)
Expand Down
82 changes: 82 additions & 0 deletions docs/api_reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -208,6 +208,88 @@ SELECT pgedge_vectorizer.set_embedding_model(
);
```

### embedding_model_status()

Report which provider and model each vectorizer's chunks were actually embedded
with, and where that disagrees with what it would use now.

```sql
SELECT * FROM pgedge_vectorizer.embedding_model_status(
source_table REGCLASS DEFAULT NULL,
source_column NAME DEFAULT NULL
);
```

**Parameters:** both optional, narrowing the result to one source table or one
column of it. With neither, every registered vectorizer is reported.

Columns:

- `source_table`, `source_column`, `chunk_table`: The vectorizer, as registered
- `effective_provider`, `effective_model`: What it would use now, inheritance
resolved
- `chunks_embedded`: Chunks with a vector. Every count below is a subset of
this one; a chunk with no vector has no model to disagree about and is
excluded throughout
- `chunks_current`: Embedded by the effective provider and model
- `chunks_other_model`: Embedded by something else. Vectors from two models are
not comparable, so these rows are effectively invisible to search
- `chunks_model_unknown`: Embedded before the extension recorded this, which is
every row on an installation that has just upgraded. Reported apart from a
mismatch because they may well be current
- `embedded_models`: The distinct `provider/model` pairs actually present,
ordered

Each row scans a chunk table, so this costs considerably more than the queue
views. A chunk table that has been dropped, or that the caller cannot read,
gives NULL counts rather than failing the whole result set.

### reembed()

Re-embed a vectorizer's chunks with the provider and model it would use now.

```sql
SELECT pgedge_vectorizer.reembed(
source_table REGCLASS,
source_column NAME,
embedding_dimension INT DEFAULT NULL
);
```

**Parameters:**

- `source_table`, `source_column`: The vectorizer to repair
- `embedding_dimension`: Dimension of the effective model. When NULL (the
default) the provider is probed for it, which is a real request

Returns: `BIGINT` - The number of chunks queued

Clears and requeues every chunk not known to have been produced by the
effective provider and model, which includes chunks with nothing recorded:
a row that cannot be shown to be current is treated as needing doing again, so
the first call on a freshly upgraded installation re-embeds the whole table.
Chunks already current are left alone.

If the effective model is a different width from the chunk table's vector
column, that distinction cannot hold: the column is altered and every chunk is
requeued, since a column cannot carry two widths. A notice says so.

Chunk rows, token counts, sparse embeddings and the BM25 statistics are
untouched either way, because none of them depends on the embedding model.

The vectorizer's pending queue rows are replaced, since anything queued was
queued before this decision was made, except rows queued for sparse work
alone: those carry no embedding to redo and are left where they are.

Unlike `set_embedding_model()`, there is no confirmation flag: this function
does what its name says. It does spend money against a metered provider, and
raises a notice with the count for that reason.

This is the supported repair for a vectorizer that drifted because
`pgedge_vectorizer.model` changed under it. `set_embedding_model()` will not do
it, because an inheriting vectorizer's effective model already is the new one,
so from that function's point of view nothing has changed.

### retry_failed()

Retry failed queue items.
Expand Down
7 changes: 7 additions & 0 deletions docs/best_practices.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,13 @@ counts survive untouched.
smaller size from providers that support shortening.
- Pin the model rather than inheriting it wherever the embeddings matter, since
an inheriting vectorizer follows `pgedge_vectorizer.model` with no guard.
- Treat a change to `pgedge_vectorizer.model` as a data migration rather than a
configuration change, because for every inheriting vectorizer that is what it
is. The sequence is: change the setting, run `embedding_model_status()` to see
which tables are now a mixture, and `reembed()` each of them, having budgeted
for embedding all of it again. Until that is done those tables hold vectors
from two models, and similarity between them is noise, so the affected rows
are effectively invisible to search rather than merely out of date.

**Performance**

Expand Down
11 changes: 11 additions & 0 deletions docs/changelog.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,6 +53,17 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
- `generate_embedding()` and `detect_embedding_dimension()` accept an optional
provider and model, so a query can be embedded with the same model as the
chunks it will be compared against.
- Chunk tables now record the provider and model that produced each embedding,
and `pgedge_vectorizer.embedding_model_status()` reports where that disagrees
with what the vectorizer would use now
([#75](https://github.com/pgEdge/pgedge-vectorizer/issues/75)). A vectorizer
that inherits follows `pgedge_vectorizer.model` as it changes, so a chunk
table can end up holding vectors from two models with nothing reporting it;
where the widths match the existing dimension check cannot see it either.
`pgedge_vectorizer.reembed()` repairs it, redoing the chunks that are not
known to be current. Rows embedded before this release have nothing recorded
and are counted separately, but `reembed()` treats them as needing redoing,
so the first call on an upgraded installation re-embeds the whole table.

### Changed

Expand Down
51 changes: 51 additions & 0 deletions docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -97,6 +97,57 @@ is not limited to changes of dimension.
setting existed. Pin the model on any vectorizer whose embeddings
matter.

Where it has already happened, it is at least visible and repairable:
`embedding_model_status()` reports what each chunk table is a mixture
of, and `reembed()` redoes the chunks that are not current. See
[Seeing what a table was embedded with](#seeing-what-a-table-was-embedded-with).

### Seeing what a table was embedded with

Every chunk records the provider and model that produced its vector, so a
disagreement between that and what the vectorizer would use now is visible
rather than something you discover through poor search results:

```sql
SELECT * FROM pgedge_vectorizer.embedding_model_status('articles'::regclass);
```

```
source_table | articles
source_column | body
effective_provider | openai
effective_model | text-embedding-3-large
chunks_embedded | 12043
chunks_current | 9945
chunks_other_model | 2098
chunks_model_unknown | 0
embedded_models | {openai/text-embedding-3-large,openai/text-embedding-3-small}
```

`chunks_other_model` is the count that matters: those rows hold vectors from a
different model, and similarity between two models' vectors is meaningless, so
they are effectively invisible to search rather than merely stale.
`chunks_model_unknown` counts rows embedded before this was recorded, which is
every row on an installation that has just upgraded; they may well be current,
so they are reported separately rather than assumed wrong.

To repair it:

```sql
SELECT pgedge_vectorizer.reembed('articles'::regclass, 'body');
```

That clears and requeues everything not known to have come from the provider
and model the vectorizer would use now, which includes the unknown rows, since
a row that cannot be shown to be current is one that needs doing again. Rows
that are already current are left alone, unless the new model is a different
width, in which case the column has to be altered and every chunk goes with it.
Chunks, token counts, sparse embeddings and the BM25 statistics are untouched
throughout, because none of them depends on the embedding model.

Both functions scan the chunk table, so give them a source table rather than
running them across every vectorizer out of habit.

## Worker Settings

These settings control the background workers that process the embedding queue, including concurrency, batch sizes, and retry behavior.
Expand Down
Loading
Loading