Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -34,7 +34,7 @@ DATA = sql/$(EXTENSION)--$(EXTVERSION).sql \
sql/$(EXTENSION)--1.0-beta3--1.0.sql

# Test configuration for pg_regress
REGRESS = setup chunking multibyte_chunking hybrid_chunking queue delete_truncate delete_truncate_pk pk_type_session max_retries vectorization multi_column maintenance edge_cases providers worker cleanup embedding pk_types stale_embeddings hybrid_test count_tokens
REGRESS = setup chunking multibyte_chunking hybrid_chunking queue delete_truncate delete_truncate_pk pk_type_session max_retries vectorization multi_column maintenance edge_cases providers worker cleanup embedding pk_types stale_embeddings hybrid_test count_tokens vectorizer_status
REGRESS_OPTS = --inputdir=test --outputdir=test

# Documentation files (if any)
Expand Down
46 changes: 46 additions & 0 deletions docs/api_reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -391,3 +391,49 @@ Count of pending items.
```sql
SELECT * FROM pgedge_vectorizer.pending_count;
```

### vectorizer_status

Embedding coverage and queue backlog for every registered vectorizer, one row
per source table and column. See
[Check Embedding Coverage](monitoring.md#check-embedding-coverage) for how to
read the numbers.

```sql
SELECT * FROM pgedge_vectorizer.vectorizer_status;
```

Columns:

- `source_table`, `source_column`, `chunk_table`: The vectorizer, as registered
- `source_rows`: Rows in the source table
- `source_rows_covered`: Source rows with at least one embedded chunk
- `source_coverage`: `source_rows_covered / source_rows`, to four decimal
places, or `NULL` for an empty source table. A value above 1 means the chunk
table holds rows for source rows that no longer exist
- `chunks_total`, `chunks_embedded`: Rows in the chunk table, and how many have
a non-NULL `embedding`
- `chunk_coverage`: `chunks_embedded / chunks_total`, to four decimal places,
or `NULL` for an empty chunk table
- `queue_pending`, `queue_processing`, `queue_failed`: Queue items for this
chunk table in each state
- `oldest_pending_age`: How long the oldest pending item has been waiting, or
`NULL` if nothing is pending
- `last_processed_at`: When an item was most recently completed. Only queue
rows that still exist are considered, so `clear_completed()` moves this
backwards

The counts scan the chunk table and the source table, so this costs
considerably more than the queue views above. If the chunk or source table has
been dropped, or the caller cannot read it, the corresponding columns are
`NULL` rather than the query failing.

The function form of the same name narrows the result to one source table, and
optionally to one column of it:

```sql
SELECT * FROM pgedge_vectorizer.vectorizer_status(
source_table REGCLASS DEFAULT NULL,
source_column NAME DEFAULT NULL
);
```
15 changes: 15 additions & 0 deletions docs/changelog.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,10 +21,25 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
particular chunk table does need bringing into line, `recreate_chunks()` on
it rewrites every row through the new path.

- Generated chunk table names are now quoted before they are looked up. The
name is built as `source_table || column || '_chunks'` and the table is
created with `%I`, so for a source table in a schema the dot ends up inside
a single identifier rather than separating a schema from a relation.
`to_regclass()` was reading that dot as qualification, which made
`recreate_chunks()` raise as though the chunk table had never been created,
and left the BM25 statistics behind when the source was truncated.

### Added

- `pgedge_vectorizer.count_tokens(text)`, which exposes the chunking engine's
token estimate so you can see why a piece of text chunked the way it did.
- `pgedge_vectorizer.vectorizer_status`, a view reporting embedding coverage
and queue backlog for each registered vectorizer, so you can tell how far
behind the embeddings are before trusting a search over them, and spot a
worker that has stalled ([#25](https://github.com/pgEdge/pgedge-vectorizer/issues/25)).
A function of the same name narrows the result to a single source table or
column. The counts scan the chunk and source tables, so this is a diagnostic
to run deliberately rather than something to poll.

## [1.1] - 2026-08-28

Expand Down
44 changes: 44 additions & 0 deletions docs/monitoring.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,3 +33,47 @@ SELECT * FROM pgedge_vectorizer.pending_count;
-- Failed items with errors
SELECT * FROM pgedge_vectorizer.failed_items;
```

## Check Embedding Coverage

Embeddings are generated asynchronously, so there is always some lag between a change to a source table and the embedding catching up with it. The `vectorizer_status` view reports, for every registered vectorizer, how much of the source is embedded and how much work is still outstanding, which is what you need in order to judge whether a search result set reflects recent changes:

```sql
SELECT * FROM pgedge_vectorizer.vectorizer_status;
```

```
source_table | articles
source_column | body
chunk_table | articles_body_chunks
source_rows | 2110
source_rows_covered | 2098
source_coverage | 0.9943
chunks_total | 12043
chunks_embedded | 11890
chunk_coverage | 0.9873
queue_pending | 153
queue_processing | 2
queue_failed | 0
oldest_pending_age | 00:04:12.882
last_processed_at | 2026-09-09 11:58:07.114+01
```

Coverage is reported two ways because they answer different questions. `source_coverage` is the fraction of source rows with at least one embedded chunk, which is the closer match to "can I trust a search over this table"; `chunk_coverage` is the fraction of individual chunks embedded, which is the better measure of how much work is left to do. A large document part-way through being embedded counts as covered on the first measure and only partly on the second.

Alongside those, `queue_pending`, `queue_processing` and `queue_failed` are the backlog for this vectorizer, `oldest_pending_age` is how long the oldest unprocessed item has been waiting, and `last_processed_at` is when an item was most recently completed. A backlog that is not shrinking and an `oldest_pending_age` that keeps growing point at a worker that is not running, or at a provider that is rejecting requests; `queue_failed` above zero is worth following up in `failed_items`.

To look at a single vectorizer rather than all of them, call the function form, optionally naming the column as well:

```sql
SELECT * FROM pgedge_vectorizer.vectorizer_status('articles'::regclass);
SELECT * FROM pgedge_vectorizer.vectorizer_status('articles'::regclass, 'body');
```

Three things are worth knowing before you put this anywhere automated.

It is not cheap. Each row counts the chunk table and the source table, so unlike the queue views above, which read only an indexed queue table, this scans your data. Treat it as a diagnostic you run when you want an answer, rather than something a dashboard polls every few seconds, and use the function form to scope it to one table where you can.

`last_processed_at` reflects only queue rows that still exist, so running `clear_completed()` will move it backwards or set it to NULL. It says when the queue last completed something it still remembers, not when the embeddings were last touched.

A `source_coverage` above 1 means the chunk table holds rows for source rows that no longer exist. That is left visible rather than clamped, because it is a real problem worth seeing: it usually means chunks were orphaned by a bulk operation that bypassed the triggers.
216 changes: 215 additions & 1 deletion sql/pgedge_vectorizer--1.1--1.2.sql
Original file line number Diff line number Diff line change
Expand Up @@ -496,7 +496,14 @@ BEGIN
END IF;

-- Verify chunk table exists
IF to_regclass(chunk_table_name) IS NULL THEN
/*
* The chunk table's name is generated as source_table || column ||
* '_chunks' and created with %I, so for a schema-qualified source the dot
* is inside a single identifier rather than separating a schema from a
* relation. Without quote_ident() the lookup splits it, finds nothing and
* raises as though the table had never been created.
*/
IF to_regclass(quote_ident(chunk_table_name)) IS NULL THEN
RAISE EXCEPTION 'Chunk table % does not exist. Use enable_vectorization() first.', chunk_table_name;
END IF;

Expand Down Expand Up @@ -622,3 +629,210 @@ $$ LANGUAGE plpgsql;

COMMENT ON FUNCTION pgedge_vectorizer.recreate_chunks IS
'Delete all chunks and recreate from source table (complete rebuild)';

---------------------------------------------------------------------------
-- Vectorizer status: how far behind the embeddings are
--
-- Embeddings are derived data generated asynchronously, so there is always
-- some lag between a source change and the embedding catching up. This
-- reports, per vectorizer, how much of the source is embedded and how much
-- work is outstanding, so that a user can tell whether a search result set
-- reflects recent changes and an operator can spot a stalled worker.
--
-- The chunk tables are named in the registry rather than joined statically,
-- so the counts have to be gathered with dynamic SQL rather than expressed
-- as a plain view. Each row costs a scan of one chunk table and a count of
-- one source table, which is considerably more than the queue views cost;
-- it is a diagnostic to run when you want an answer, not something to put
-- on a dashboard refreshing every second.
---------------------------------------------------------------------------

CREATE FUNCTION pgedge_vectorizer.vectorizer_status(
p_source_table REGCLASS DEFAULT NULL,
p_source_column NAME DEFAULT NULL
) RETURNS TABLE (
source_table TEXT,
source_column NAME,
chunk_table TEXT,
source_rows BIGINT,
source_rows_covered BIGINT,
source_coverage NUMERIC,
chunks_total BIGINT,
chunks_embedded BIGINT,
chunk_coverage NUMERIC,
queue_pending BIGINT,
queue_processing BIGINT,
queue_failed BIGINT,
oldest_pending_age INTERVAL,
last_processed_at TIMESTAMPTZ
) AS $$
DECLARE
v RECORD;
src_oid OID;
chunk_oid OID;
BEGIN
FOR v IN
SELECT r.source_table, r.source_column, r.chunk_table
FROM pgedge_vectorizer.vectorizers r
WHERE (p_source_column IS NULL OR r.source_column = p_source_column)
ORDER BY r.source_table, r.source_column
LOOP
/*
* to_regclass() does not return NULL for every name it cannot
* resolve: given a qualified name whose schema the caller has no
* USAGE on, it raises insufficient_privilege instead, so resolving
* the source table has to be guarded. Without the guard a caller
* holding rights on one vectorizer and not on another's schema gets
* an error for the whole result set rather than the rows it can
* see, and the has_table_privilege() check further down never runs.
*
* The narrowing by p_source_table is applied here rather than in the
* query above for the same reason: resolving every registry row in
* the WHERE clause would raise on a vectorizer the caller cannot see
* even when it asked about a different table.
*/
BEGIN
src_oid := to_regclass(v.source_table);
EXCEPTION WHEN insufficient_privilege THEN
src_oid := NULL;
END;

CONTINUE WHEN p_source_table IS NOT NULL
AND src_oid IS DISTINCT FROM p_source_table::OID;

source_table := v.source_table;
source_column := v.source_column;
chunk_table := v.chunk_table;

source_rows := NULL;
source_rows_covered := NULL;
source_coverage := NULL;
chunks_total := NULL;
chunks_embedded := NULL;
chunk_coverage := NULL;

/*
* Queue figures first: they come from a table the caller can always
* read, and they are the half that stays meaningful even when the
* chunk or source table has been dropped from under the registry.
*
* last_processed_at only reflects queue rows that still exist, so
* clear_completed() will move it backwards. That is a property of
* the queue, not of the embeddings, and is documented as such.
*/
SELECT count(*) FILTER (WHERE q.status = 'pending'),
count(*) FILTER (WHERE q.status = 'processing'),
count(*) FILTER (WHERE q.status = 'failed'),
NOW() - min(q.created_at) FILTER (WHERE q.status = 'pending'),
max(q.processed_at)
INTO queue_pending, queue_processing, queue_failed,
oldest_pending_age, last_processed_at
FROM pgedge_vectorizer.queue q
WHERE q.chunk_table = v.chunk_table;

/*
* The counts below read user tables, and the function runs as the
* caller. A table that has been dropped, or that the caller cannot
* read, leaves the corresponding columns NULL rather than failing
* the whole result set: one inaccessible vectorizer should not make
* the view useless for every other one.
*/
/*
* The chunk table's name is generated as source_table || column ||
* '_chunks' and created with %I, so for a schema-qualified source the
* dot ends up inside a single identifier rather than separating a
* schema from a relation. quote_ident() keeps to_regclass() from
* splitting it, matching how quote_identifier() is used for the same
* name in bm25.c. The source table's name, by contrast, comes from
* regclass output and is a genuine qualified reference, so it is
* looked up as it stands.
*/
chunk_oid := to_regclass(quote_ident(v.chunk_table));
IF chunk_oid IS NOT NULL
AND has_table_privilege(chunk_oid, 'SELECT') THEN
EXECUTE format(
'SELECT count(*),
count(*) FILTER (WHERE embedding IS NOT NULL),
count(DISTINCT source_id)
FILTER (WHERE embedding IS NOT NULL)
FROM %s', chunk_oid::REGCLASS)
INTO chunks_total, chunks_embedded, source_rows_covered;

IF chunks_total > 0 THEN
chunk_coverage := round(chunks_embedded::NUMERIC
/ chunks_total, 4);
END IF;
END IF;

IF src_oid IS NOT NULL
AND has_table_privilege(src_oid, 'SELECT') THEN
EXECUTE format('SELECT count(*) FROM %s', src_oid::REGCLASS)
INTO source_rows;

/*
* A coverage above 1 means the chunk table holds rows for source
* rows that are gone, which is worth seeing rather than hiding
* behind a clamp to 1.
*/
IF source_rows > 0 AND source_rows_covered IS NOT NULL THEN
source_coverage := round(source_rows_covered::NUMERIC
/ source_rows, 4);
END IF;
END IF;

RETURN NEXT;
END LOOP;
END;
$$ LANGUAGE plpgsql;

COMMENT ON FUNCTION pgedge_vectorizer.vectorizer_status IS
'Embedding coverage and queue backlog for the registered vectorizers, '
'optionally narrowed to one source table or one source column. Scans the '
'chunk and source tables, so it costs considerably more than the queue views';

CREATE VIEW pgedge_vectorizer.vectorizer_status AS
SELECT * FROM pgedge_vectorizer.vectorizer_status(NULL, NULL);

COMMENT ON VIEW pgedge_vectorizer.vectorizer_status IS
'Embedding coverage and queue backlog for every registered vectorizer';

---------------------------------------------------------------------------
-- Redefine the TRUNCATE trigger function for the same identifier reason
--
-- It looked up the BM25 statistics table without quoting, so for a
-- schema-qualified source the dot in the generated name was read as
-- qualification, the lookup came back NULL and truncating the source left
-- the corpus statistics behind for chunks that no longer existed. The
-- trigger itself is unchanged, so only the function needs replacing.
---------------------------------------------------------------------------

CREATE OR REPLACE FUNCTION pgedge_vectorizer.vectorization_truncate_trigger()
RETURNS TRIGGER AS $$
DECLARE
chunk_table TEXT;
BEGIN
chunk_table := TG_ARGV[0];

EXECUTE format(
'DELETE FROM pgedge_vectorizer.queue
WHERE chunk_table = %L AND status IN (''pending'', ''failed'')',
chunk_table);

EXECUTE format('TRUNCATE TABLE %I', chunk_table);

/*
* Both names are single identifiers, generated from the source table and
* column and created with %I, so a schema-qualified source leaves a dot
* inside the identifier. quote_ident() stops to_regclass() reading that
* dot as qualification and returning NULL for a table that is there.
*/
IF to_regclass(quote_ident(chunk_table || '_idf_stats')) IS NOT NULL THEN
EXECUTE format('TRUNCATE TABLE %I', chunk_table || '_idf_stats');
END IF;

RETURN NULL;
END;
$$ LANGUAGE plpgsql;

COMMENT ON FUNCTION pgedge_vectorizer.vectorization_truncate_trigger IS
'Statement-level AFTER TRUNCATE trigger emptying the chunk table, its queue entries and its BM25 statistics';
Loading
Loading