mirror of
https://github.com/willmiao/ComfyUI-Lora-Manager.git
synced 2026-09-21 11:11:26 -03:00
feat(download): fill model metadata from the source API on download
A ModelScope or Hugging Face download landed as a bare filename, hash and
source link; the model card stayed empty until the user ran "Enrich
Metadata with AI" by hand. But everything that makes a CivitAI download
useful — the display name, the description, the tags, the trigger words,
the example images, the preview — is already published by those sites'
public APIs, so asking for it at download time is deterministic work, not
model work.
Add `py/services/model_sources/hydration.py`, called by
`_save_source_metadata()` once the sidecar exists and the file is in the
scanner cache. It fetches the model card plus the site's card extras and
hands them to the same `PostProcessor` the AI skill uses, with an empty
`llm_output`, so the two paths cannot drift apart. What lands:
* `model_name` from the site's own display name (ModelScope's `Name`), so
the card stops showing the local filename — written only while the value
still equals the file stem, since once a user renames a model that
choice is theirs to keep
* `civitai.name` from the matched version's label (`showName`), which the
card renders as the version chip
* `civitai.description` / `modelDescription` from the author summary plus
the README as HTML
* `civitai.images` / `preview_url` from the per-file example images
* `civitai.trainedWords` from the per-file trigger words
* `base_model`, `tags` and `usage_tips` as before
Provenance stays honest: the pass records
`metadata_source = "source:<platform>"` rather than the skill's
`agent:enrich_hf_metadata`, and — because no provider ran — it no longer
stamps `llm_enriched_at`; that stamp is now conditional on the LLM
actually answering, which is what the field means. The five hand-rolled
`civitai` dict merges in the post-processor collapse into one
`_merge_civitai()` helper.
Two guards keep it safe. Only a model whose stored
`source_platform`/`source_url` match the repository being downloaded is
updated, so a local file that merely shares a name never receives another
model's card; and a file already on disk is topped up too, which
back-fills models downloaded before this existed. READMEs and detail
payloads describe the repository rather than the file, so a short-lived
process-wide `ModelSourceCache` (300 s, 32 entries) keeps a batch over one
repository to two HTTP requests. Every failure is logged and swallowed:
hydration can never fail a download.
Fix the hash policy while here. `_save_source_metadata()` went straight to
`MetadataManager.create_default_metadata()`, bypassing the per-type
factory on the owning scanner, so a checkpoint paid a full SHA256 inside
the download request — `CheckpointScanner`/`OtherScanner` deliberately
record `hash_status="pending"` with an empty `sha256` for their multi-GB
files. Metadata is now created through `scanner._create_default_metadata()`.
Hydration copes with the empty hash: `_matching_versions()` falls back to
the repository basename, which is exactly what the download just wrote.
Report both post-transfer stages, which advance no byte counter and so
read as a stall: the bar sat at 100% showing `0 B/s` for the seconds spent
hashing and fetching. `_report_phase()` broadcasts
`{"status": "metadata", "stage": "indexing" | "source", "platform": ...}`,
and `LoadingManager` names the stage in the status line (keeping the batch
position), retitles the item line, replaces the dead speed figure and runs
a sheen over the bar. `stage`/`platform` are machine-readable; the wording
is localised in the frontend.
Finally, `modelscope.ai` is its own catalogue rather than an alias of
`modelscope.cn` — `referall13/EM1` exists only on `.ai` and
`jj3550945163/Krea-2-LORA` only on `.cn` — so its URLs were rejected with
"Invalid model URL format". Register it as `ModelScopeIntlSource`
(`platform="modelscope-ai"`, `msai:` group prefix, its own default
download directory) and derive every URL either deployment builds from a
per-class `base_url`. `modelscope.com` stays an alias of `.cn`, which is
what it redirects to. The frontend source table, the link dialog hints and
the docs mirror the split.
Verified against the live APIs: both reported `.ai` repositories list
their files, read their READMEs and yield name / version / base model /
trigger words / example images. Backend 3092 passed; frontend 1259 JS +
91 Vue passed. The nine locales carry the new progress copy in the next
commit.
This commit is contained in:
+53
-1
@@ -71,9 +71,18 @@ Enriches models linked to an external model site with metadata extracted by an L
|
||||
| Platform | Link | AI enrichment | Direct download |
|
||||
| --- | --- | --- | --- |
|
||||
| Hugging Face | yes | yes | yes |
|
||||
| ModelScope | yes | yes | yes |
|
||||
| ModelScope (`modelscope.cn`) | yes | yes | yes |
|
||||
| ModelScope International (`modelscope.ai`) | yes | yes | yes |
|
||||
| TensorArt | yes | no (see below) | no |
|
||||
|
||||
`modelscope.cn` and `modelscope.ai` are **separate catalogues, not mirrors** — a
|
||||
repository published on one is routinely absent from the other — so each is
|
||||
registered as its own source (`ModelScopeSource` / `ModelScopeIntlSource` in
|
||||
`py/services/model_sources/modelscope.py`). The host therefore decides which
|
||||
API and CDN a model resolves against, and the two deployments get separate
|
||||
version groups (`ms:` / `msai:`) and default download directories. Keep the two
|
||||
tables in `modelSourceHelpers.js` and `registry.py` in step when adding a site.
|
||||
|
||||
TensorArt is link-only: `tensor.art` sits behind a Cloudflare managed challenge and its internal API requires session authorization, so the backend cannot read its model pages. Linking still stores the canonical page URL and the "View on TensorArt" link works.
|
||||
|
||||
**What it does**:
|
||||
@@ -133,7 +142,9 @@ gaps the LLM leaves behind:
|
||||
|
||||
| Field | Deterministic source | LLM role |
|
||||
| --- | --- | --- |
|
||||
| `model_name` | site display name (`Name`), written only while the value is still the file stem | — |
|
||||
| `modelDescription` | author summary + README as HTML | — |
|
||||
| `civitai.name` | the matched version's label (`modelVersion.showName`) | — |
|
||||
| `civitai.images` | site example images, then README images | — |
|
||||
| `preview_url` | first available example image | may propose one from the README |
|
||||
| `tags` | site-curated tags, always merged in | proposes additional content tags |
|
||||
@@ -147,6 +158,47 @@ Models with no source, an unknown source, or a source without model-card access
|
||||
|
||||
**Model types**: LoRA, Checkpoint, Embedding
|
||||
|
||||
### Download-time hydration
|
||||
|
||||
The same deterministic mapping runs automatically when a model is downloaded
|
||||
from a model source, so a ModelScope or Hugging Face download lands with the
|
||||
populated card a CivitAI download produces instead of a bare filename and
|
||||
hash. Nothing needs to be triggered by hand and no provider is called.
|
||||
|
||||
`py/services/model_sources/hydration.py` owns this path:
|
||||
|
||||
* `_save_source_metadata()` in `py/routes/handlers/model_source_handlers.py`
|
||||
creates the sidecar (hash, source link, scanner-cache entry) and then calls
|
||||
`hydrate_from_source()`. It also runs for a file that was already on disk, so
|
||||
models downloaded before this existed get topped up on the next attempt.
|
||||
* Metadata is created through the **owning scanner**
|
||||
(`scanner._create_default_metadata()`) rather than
|
||||
`MetadataManager.create_default_metadata()`, so the per-type lazy-hash rule
|
||||
applies: `CheckpointScanner` and `OtherScanner` store
|
||||
`hash_status="pending"` with an empty `sha256` for their multi-GB files, and
|
||||
the generic helper would read a 10 GB checkpoint end to end inside the
|
||||
download request. Hydration copes with the empty hash — `_matching_versions()`
|
||||
falls back to the repository basename, which the download just wrote.
|
||||
* Hydration reuses `PostProcessor` with an empty `llm_output`, so the two paths
|
||||
cannot drift apart. It reports `metadata_source = "source:<platform>"` rather
|
||||
than the skill's `agent:enrich_hf_metadata`, and — because no provider ran —
|
||||
it does not stamp `llm_enriched_at`.
|
||||
* `model_name` is only written while it still equals the file stem: once a user
|
||||
renames a model, that choice is kept.
|
||||
* Only a model whose stored `source_platform`/`source_url` match the repository
|
||||
being downloaded is updated; a local file that merely shares a name must not
|
||||
receive another model's card.
|
||||
* The README and repository payload describe the *repository*, so a short-lived
|
||||
process-wide `ModelSourceCache` (`shared_source_cache`, 300 s, 32 entries)
|
||||
keeps a batch over one repository to two HTTP requests.
|
||||
* Every failure — unreachable site, changed payload shape, broken post-processor
|
||||
— is logged and swallowed. Metadata hydration can never fail a download.
|
||||
* Neither stage advances the byte counter, so both are announced to the
|
||||
progress UI (`_report_phase()` → `{"status": "metadata", "stage": ...}`) as
|
||||
they start. Without that the bar sits at 100% reporting `0 B/s` for several
|
||||
seconds and the download looks stuck. `stage` and `platform` are
|
||||
machine-readable; the wording is localised in `LoadingManager`.
|
||||
|
||||
## Adding a New Skill
|
||||
|
||||
### 1. Create the skill directory
|
||||
|
||||
Reference in New Issue
Block a user