Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
33 commits
Select commit Hold shift + click to select a range
6e0ef98
Add modular analysis backend config and distribution schema
sanghoonio Mar 27, 2026
f0ecd82
Fix bedset stats aggregation: skip non-numeric columns
sanghoonio Apr 3, 2026
0558584
Cache get_stats() with a TTL to fix uncached COUNT on hot API paths
nsheff Jul 13, 2026
1ceb219
Batch neighbour metadata fetch to eliminate N+1 in get_neighbours
nsheff Jul 14, 2026
1aa8912
Improved bedset search
khoroshevskyi Aug 1, 2026
5e7e855
improvement efficiency of bedbase detailed stats
khoroshevskyi Aug 1, 2026
c9fd9a0
Fixed detailed stats fetching bug
khoroshevskyi Aug 1, 2026
6674386
linting
khoroshevskyi Aug 3, 2026
a9a1fbc
Add bed_snapshots table and export index models
nsheff Aug 3, 2026
2c02f4f
Bump version to 0.14.13
nsheff Aug 3, 2026
9b5685b
Merge branch 'efficiency_improvements' into feat/metadata-export
khoroshevskyi Aug 3, 2026
05f1467
Polishing downloads / snapshots endpoints
khoroshevskyi Aug 4, 2026
34068d3
Fixed incorrect search count
khoroshevskyi Aug 15, 2026
5ee9a48
Added some important indexes
khoroshevskyi Aug 15, 2026
8bbc2d5
Added all necessury indexes to the tables
khoroshevskyi Aug 15, 2026
692f3ff
Added alembic for database migrations
khoroshevskyi Aug 15, 2026
340d56a
deleted unused arguemnt
khoroshevskyi Aug 15, 2026
37d87f5
Small tweaks
khoroshevskyi Aug 15, 2026
c3fd35f
Linting and test fix
khoroshevskyi Aug 16, 2026
0b48d4e
Removed phc from bedbase
khoroshevskyi Aug 16, 2026
9872e9b
removed leftovers of pephubclient
khoroshevskyi Aug 16, 2026
364a960
Merge pull request #123 from databio/alembic
khoroshevskyi Aug 16, 2026
e7e1e1f
Merge pull request #120 from databio/efficiency_improvements
khoroshevskyi Aug 16, 2026
7419059
Merge branch 'dev' into modular-backend-schema
khoroshevskyi Aug 16, 2026
a9c1367
updated gtars distribution new schema
khoroshevskyi Aug 16, 2026
96c05bd
Fixed config file initialization
khoroshevskyi Aug 16, 2026
de184b0
Added alembic migration for distribution
khoroshevskyi Aug 16, 2026
a2b1aea
Merge pull request #112 from databio/modular-backend-schema
khoroshevskyi Aug 16, 2026
725cc92
Updated changelog and version
khoroshevskyi Aug 17, 2026
43d2441
Merge remote-tracking branch 'origin/dev' into dev
khoroshevskyi Aug 17, 2026
565e1dc
added Analysis files functionality
khoroshevskyi Aug 18, 2026
85dfa8b
Added back alembic file
khoroshevskyi Aug 18, 2026
ad87474
Merge pull request #125 from databio/analysis_files
khoroshevskyi Aug 18, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .github/pull_request_template.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,5 +3,6 @@


## TODO:
- [ ] ❗ If this PR includes a new database schema migration, following steps are completed: (README)[README.md]
- [ ] Version of pepdbagent updated in `__version__.py` file
- [ ] Changelog updated
1 change: 1 addition & 0 deletions .github/workflows/run-pytest.yml
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,7 @@ jobs:
python-version: ["3.10", "3.13"]
os: [ubuntu-latest] # can't use macOS when using service containers or container jobs
runs-on: ${{ matrix.os }}

services:
postgres:
image: postgres
Expand Down
57 changes: 54 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,11 +48,62 @@ from bbconf import BedBaseAgent
agent = BedBaseAgent(config="config.yaml")

# Access submodules
agent.bed # BED file operations
agent.bedset # BED set operations
agent.objects # Generic object/file operations
agent.bed # BED file operations
agent.bedset # BED set operations
agent.objects # Generic object/file operations

# Get platform statistics
stats = agent.get_stats()
print(stats.bedfiles_number, stats.bedsets_number)
```

## Database migrations

`bbconf` uses [Alembic](https://alembic.sqlalchemy.org/) to version the database
schema. The migration scripts live in `bbconf/alembic`, and `alembic.ini` (repo
root) is used for local CLI work. The first (baseline) revision is
`8b0b706d0827`; it reproduces exactly the schema that `Base.metadata.create_all()`
builds, including the `pg_trgm` extension and the trigram / partial / expression
indexes.

To update schema for desirable database, use different database url in `alembic.ini`,
otherwise run test database

### Creating a new revision

After changing the models in `bbconf/db_utils.py`:

```bash
alembic revision --autogenerate -m "Describe your change"
```

Review the generated file. Alembic cannot autogenerate a few constructs used by
bbconf — the `pg_trgm` extension and expression-based indexes may need a manual
`op.execute(...)` — so always check the diff before committing.

### Applying migrations

```bash
alembic upgrade head # upgrade to the latest revision
alembic downgrade -1 # roll back one revision
alembic current # show the DB's current revision
```

### Running migrations automatically

To upgrade the database to `head` automatically when `bbconf` starts, set
`run_migrations: true` under the `database` section of the config file:

```yaml
database:
host: localhost
port: 5432
user: postgres
password: docker
database: bedbase
run_migrations: true
```

> **Note:** enable this only after the database has been stamped/upgraded to a
> known revision. Turning it on against an un-stamped existing database will fail
> on startup, because the baseline revision creates tables that already exist.
128 changes: 128 additions & 0 deletions alembic.ini
Original file line number Diff line number Diff line change
@@ -0,0 +1,128 @@
# A generic, single database configuration.
#
# This file is used only for local development / CLI work
# (e.g. `alembic revision --autogenerate`, `alembic upgrade head`).
# At runtime, bbconf builds the Alembic config programmatically in
# `BaseEngine.run_db_migration()` and does NOT read this file.

[alembic]
# path to migration scripts
# Use forward slashes (/) also on windows to provide an os agnostic path
script_location = ./bbconf/alembic

# template used to generate migration file names; The default value is %%(rev)s_%%(slug)s
# Uncomment the line below if you want the files to be prepended with date and time
# see https://alembic.sqlalchemy.org/en/latest/tutorial.html#editing-the-ini-file
# for all available tokens
# file_template = %%(year)d_%%(month).2d_%%(day).2d_%%(hour).2d%%(minute).2d-%%(rev)s_%%(slug)s

# sys.path path, will be prepended to sys.path if present.
# defaults to the current working directory.
prepend_sys_path = .

# timezone to use when rendering the date within the migration file
# as well as the filename.
# If specified, requires the python>=3.9 or backports.zoneinfo library and tzdata library.
# Any required deps can installed by adding `alembic[tz]` to the pip requirements
# string value is passed to ZoneInfo()
# leave blank for localtime
# timezone =

# max length of characters to apply to the "slug" field
# truncate_slug_length = 40

# set to 'true' to run the environment during
# the 'revision' command, regardless of autogenerate
# revision_environment = false

# set to 'true' to allow .pyc and .pyo files without
# a source .py file to be detected as revisions in the
# versions/ directory
# sourceless = false

# version location specification; This defaults
# to alembic/versions. When using multiple version
# directories, initial revisions must be specified with --version-path.
# The path separator used here should be the separator specified by "version_path_separator" below.
# version_locations = %(here)s/bar:%(here)s/bat:alembic/versions

# version path separator; As mentioned above, this is the character used to split
# version_locations. The default within new alembic.ini files is "os", which uses os.pathsep.
# If this key is omitted entirely, it falls back to the legacy behavior of splitting on spaces and/or commas.
# Valid values for version_path_separator are:
#
# version_path_separator = :
# version_path_separator = ;
# version_path_separator = space
# version_path_separator = newline
#
# Use os.pathsep. Default configuration used for new projects.
version_path_separator = os

# set to 'true' to search source files recursively
# in each "version_locations" directory
# new in Alembic version 1.10
# recursive_version_locations = false

# the output encoding used when revision files
# are written from script.py.mako
# output_encoding = utf-8

# Local development connection string. Override with `-x` or edit as needed.
# Runtime migrations use the URL built from the bbconf config instead.

### !!!! Change this code to desirable database!!!!
sqlalchemy.url = postgresql+psycopg://postgres:docker@localhost:5432/bedbase


[post_write_hooks]
# post_write_hooks defines scripts or Python functions that are run
# on newly generated revision scripts. See the documentation for further
# detail and examples

# format using "black" - use the console_scripts runner, against the "black" entrypoint
# hooks = black
# black.type = console_scripts
# black.entrypoint = black
# black.options = -l 79 REVISION_SCRIPT_FILENAME

# lint with attempts to fix using "ruff" - use the exec runner, execute a binary
# hooks = ruff
# ruff.type = exec
# ruff.executable = %(here)s/.venv/bin/ruff
# ruff.options = check --fix REVISION_SCRIPT_FILENAME

# Logging configuration
[loggers]
keys = root,sqlalchemy,alembic

[handlers]
keys = console

[formatters]
keys = generic

[logger_root]
level = WARNING
handlers = console
qualname =

[logger_sqlalchemy]
level = WARNING
handlers =
qualname = sqlalchemy.engine

[logger_alembic]
level = INFO
handlers =
qualname = alembic

[handler_console]
class = StreamHandler
args = (sys.stderr,)
level = NOTSET
formatter = generic

[formatter_generic]
format = %(levelname)-5.5s [%(name)s] %(message)s
datefmt = %H:%M:%S
1 change: 1 addition & 0 deletions bbconf/alembic/README
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Generic single-database configuration.
Empty file added bbconf/alembic/__init__.py
Empty file.
77 changes: 77 additions & 0 deletions bbconf/alembic/env.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,77 @@
from logging.config import fileConfig

from alembic import context
from sqlalchemy import engine_from_config, pool

# this is the Alembic Config object, which provides
# access to the values within the .ini file in use.
config = context.config

# Interpret the config file for Python logging.
# This line sets up loggers basically.
if config.config_file_name is not None:
fileConfig(config.config_file_name)

# add your model's MetaData object here
# for 'autogenerate' support.
# Importing from bbconf.db_utils also registers the custom @compiles types
# (BIGSERIAL, JSON->JSONB, ARRAY) and the pg_trgm extension DDL event, so the
# metadata compiles to exactly the same DDL that create_all() produces.
from bbconf.db_utils import Base

target_metadata = Base.metadata

# other values from the config, defined by the needs of env.py,
# can be acquired:
# my_important_option = config.get_main_option("my_important_option")
# ... etc.


def run_migrations_offline() -> None:
"""Run migrations in 'offline' mode.

This configures the context with just a URL
and not an Engine, though an Engine is acceptable
here as well. By skipping the Engine creation
we don't even need a DBAPI to be available.

Calls to context.execute() here emit the given string to the
script output.

"""
url = config.get_main_option("sqlalchemy.url")
context.configure(
url=url,
target_metadata=target_metadata,
literal_binds=True,
dialect_opts={"paramstyle": "named"},
)

with context.begin_transaction():
context.run_migrations()


def run_migrations_online() -> None:
"""Run migrations in 'online' mode.

In this scenario we need to create an Engine
and associate a connection with the context.

"""
connectable = engine_from_config(
config.get_section(config.config_ini_section, {}),
prefix="sqlalchemy.",
poolclass=pool.NullPool,
)

with connectable.connect() as connection:
context.configure(connection=connection, target_metadata=target_metadata)

with context.begin_transaction():
context.run_migrations()


if context.is_offline_mode():
run_migrations_offline()
else:
run_migrations_online()
28 changes: 28 additions & 0 deletions bbconf/alembic/script.py.mako
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
"""${message}

Revision ID: ${up_revision}
Revises: ${down_revision | comma,n}
Create Date: ${create_date}

"""
from typing import Sequence, Union

import sqlalchemy as sa
from alembic import op
${imports if imports else ""}

# revision identifiers, used by Alembic.
revision: str = ${repr(up_revision)}
down_revision: Union[str, None] = ${repr(down_revision)}
branch_labels: Union[str, Sequence[str], None] = ${repr(branch_labels)}
depends_on: Union[str, Sequence[str], None] = ${repr(depends_on)}


def upgrade() -> None:
"""Upgrade schema."""
${upgrades if upgrades else "pass"}


def downgrade() -> None:
"""Downgrade schema."""
${downgrades if downgrades else "pass"}
Original file line number Diff line number Diff line change
@@ -0,0 +1,60 @@
"""Added genomic distribution json plots

Revision ID: 845d978eac7d
Revises: 8b0b706d0827
Create Date: 2026-08-16 18:02:03.058352

"""

from typing import Sequence, Union

import sqlalchemy as sa
from alembic import op
from sqlalchemy.dialects import postgresql

# revision identifiers, used by Alembic.
revision: str = "845d978eac7d"
down_revision: Union[str, None] = "8b0b706d0827"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None


def upgrade() -> None:
"""Upgrade schema."""

op.drop_column("bed", "pephub")
op.add_column(
"bed_stats",
sa.Column(
"distributions",
postgresql.JSONB(astext_type=sa.Text()),
nullable=True,
comment="Full distribution arrays from gtars genomicdist (JSONB)",
),
)
op.add_column(
"bedsets",
sa.Column(
"bedset_stats",
postgresql.JSONB(astext_type=sa.Text()),
nullable=True,
comment="Pre-aggregated distribution statistics from gtars (JSONB)",
),
)


def downgrade() -> None:
"""Downgrade schema."""
# ### commands auto generated by Alembic - please adjust! ###
op.drop_column("bedsets", "bedset_stats")
op.drop_column("bed_stats", "distributions")
op.add_column(
"bed",
sa.Column(
"pephub",
sa.BOOLEAN(),
autoincrement=False,
nullable=False,
comment="Whether sample was added to pephub",
),
)
Loading