feat(metadata): add OpenModelDB metadata provider and model source for upscalers

Add OpenModelDB (openmodeldb.info) as a metadata and download source for
the existing upscaler model type.

Metadata:
- New OpenModelDBClient: fetches the site's bulk JSON dumps, caches them
  on disk (24h TTL + ETag revalidation), and builds a local sha256 index
- New OpenModelDBModelMetadataProvider adapts catalogue entries to the
  CivitAI-shaped version dict contract; registered in the fallback chain
  behind the enable_openmodeldb_api setting (default on), gated to the
  upscaler sub-type so other model types never trigger the dump download
- Persisted provenance uses metadata_source "openmodeldb" plus a nested
  openmodeldb block (page URL, architecture, scale, license)

Images: paired-image LR/SR URLs are ephemeral imgdiff.net sessions, so
displayable images come from the site-hosted auto-generated thumbnails
(model-level cover leads images[], per-image thumbs for the rest); the
original comparison URL is kept in meta.comparisonUrl.

Downloads:
- New OpenModelDBSource (flat model ids, omdb: group prefix) with
  resource filename derivation that recovers names hidden mid-path
  (mediafire) or synthesizes {id}.{type} for folder links
- HTML-gateway mirrors (mediafire/mega/drive) are rejected with a clear
  manual-download hint instead of silently saving an HTML page as .pth
- ModelSource base gains is_valid_source_id / default_subdir_parts /
  resolve_download_url hooks so flat-id sources need no platform branches

UI: "View on OpenModelDB" link in the model modal (downloaded and
hash-enriched models), settings toggle next to the CivArchive one.
This commit is contained in:
Will Miao
2026-10-03 21:13:07 +08:00
parent 515469054c
commit 35f1ced41a
39 changed files with 2787 additions and 37 deletions
+3
View File
@@ -74,6 +74,7 @@ Enriches models linked to an external model site with metadata extracted by an L
| ModelScope (`modelscope.cn`) | yes | yes | yes |
| ModelScope International (`modelscope.ai`) | yes | yes | yes |
| TensorArt | yes | no (see below) | no |
| OpenModelDB | yes | yes (card data from the catalogue, no README) | yes |
`modelscope.cn` and `modelscope.ai` are **separate catalogues, not mirrors** — a
repository published on one is routinely absent from the other — so each is
@@ -85,6 +86,8 @@ tables in `modelSourceHelpers.js` and `registry.py` in step when adding a site.
TensorArt is link-only: `tensor.art` sits behind a Cloudflare managed challenge and its internal API requires session authorization, so the backend cannot read its model pages. Linking still stores the canonical page URL and the "View on TensorArt" link works.
OpenModelDB is the upscaler catalogue: model ids are flat tokens (no `owner/name`), there are no revisions and no README — `fetch_model_card_context()` reads everything (description, license, tags, example images) from the disk-cached bulk catalogue in `py/services/openmodeldb_client.py`. Only PyTorch resources (`.pth`/`.safetensors`) are downloadable, and only via mirrors that serve raw bytes: HTML-gateway hosts (`mediafire.com`, `mega.nz`, `drive.google.com`) are skipped in favour of a direct mirror, and a model with only gateway mirrors reports a manual-download hint instead of a file list. Filenames are derived from URL path segments (mediafire buries them mid-path) or synthesized as `{model_id}.{type}` for folder links.
**What it does**:
1. Reads the model's `.metadata.json` to get the source (`source_platform` + `source_url`, or the legacy `hf_url`)
2. Fetches the model card through the provider in `py/services/model_sources/` — the README via `fetch_model_card()`, plus any extras the site keeps outside it via `fetch_model_card_context()`
+20
View File
@@ -299,9 +299,29 @@ The `metadata_source` field indicates which provider last updated the metadata:
|-------|--------|
| `"civitai_api"` | Civitai API |
| `"civarchive"` | CivArchive API |
| `"openmodeldb"` | OpenModelDB catalogue (upscaler models only; hash-matched) |
| `"archive_db"` | Metadata Archive Database |
| `null` | No external source (user-defined only) |
When `metadata_source` is `"openmodeldb"`, the `civitai` payload is a
CivitAI-shaped version dict synthesized from the OpenModelDB catalogue entry
(no numeric `id`/`modelId`), and the OpenModelDB-native details live in its
`openmodeldb` block (`id`, `url`, `authors`, `architecture`,
`architectureName`, `scale`, `license`, `date`).
In that payload, `images[].url` is always a displayable asset: paired
comparisons use the site-hosted thumbnail because the `LR`/`SR` originals
are ephemeral imgdiff viewer sessions that 404 outside them (the original
viewer link is kept as `images[].meta.comparisonUrl` for reference), and the
model-level thumbnail leads the list since the card preview derives from
`images[0]`. `files[].name` is derived from any URL path segment with a
model extension (mediafire-style mirrors bury it mid-path) or synthesized as
`{model_id}.{type}` for folder links.
Models downloaded from OpenModelDB additionally carry `source_platform:
"openmodeldb"` and `source_url` (the model page URL); their download-time
hydration is recorded as `metadata_source: "source:openmodeldb"`.
---
## Auto-Update Behavior