feat(download): fill model metadata from the source API on download

A ModelScope or Hugging Face download landed as a bare filename, hash and
source link; the model card stayed empty until the user ran "Enrich
Metadata with AI" by hand. But everything that makes a CivitAI download
useful — the display name, the description, the tags, the trigger words,
the example images, the preview — is already published by those sites'
public APIs, so asking for it at download time is deterministic work, not
model work.

Add `py/services/model_sources/hydration.py`, called by
`_save_source_metadata()` once the sidecar exists and the file is in the
scanner cache. It fetches the model card plus the site's card extras and
hands them to the same `PostProcessor` the AI skill uses, with an empty
`llm_output`, so the two paths cannot drift apart. What lands:

* `model_name` from the site's own display name (ModelScope's `Name`), so
  the card stops showing the local filename — written only while the value
  still equals the file stem, since once a user renames a model that
  choice is theirs to keep
* `civitai.name` from the matched version's label (`showName`), which the
  card renders as the version chip
* `civitai.description` / `modelDescription` from the author summary plus
  the README as HTML
* `civitai.images` / `preview_url` from the per-file example images
* `civitai.trainedWords` from the per-file trigger words
* `base_model`, `tags` and `usage_tips` as before

Provenance stays honest: the pass records
`metadata_source = "source:<platform>"` rather than the skill's
`agent:enrich_hf_metadata`, and — because no provider ran — it no longer
stamps `llm_enriched_at`; that stamp is now conditional on the LLM
actually answering, which is what the field means. The five hand-rolled
`civitai` dict merges in the post-processor collapse into one
`_merge_civitai()` helper.

Two guards keep it safe. Only a model whose stored
`source_platform`/`source_url` match the repository being downloaded is
updated, so a local file that merely shares a name never receives another
model's card; and a file already on disk is topped up too, which
back-fills models downloaded before this existed. READMEs and detail
payloads describe the repository rather than the file, so a short-lived
process-wide `ModelSourceCache` (300 s, 32 entries) keeps a batch over one
repository to two HTTP requests. Every failure is logged and swallowed:
hydration can never fail a download.

Fix the hash policy while here. `_save_source_metadata()` went straight to
`MetadataManager.create_default_metadata()`, bypassing the per-type
factory on the owning scanner, so a checkpoint paid a full SHA256 inside
the download request — `CheckpointScanner`/`OtherScanner` deliberately
record `hash_status="pending"` with an empty `sha256` for their multi-GB
files. Metadata is now created through `scanner._create_default_metadata()`.
Hydration copes with the empty hash: `_matching_versions()` falls back to
the repository basename, which is exactly what the download just wrote.

Report both post-transfer stages, which advance no byte counter and so
read as a stall: the bar sat at 100% showing `0 B/s` for the seconds spent
hashing and fetching. `_report_phase()` broadcasts
`{"status": "metadata", "stage": "indexing" | "source", "platform": ...}`,
and `LoadingManager` names the stage in the status line (keeping the batch
position), retitles the item line, replaces the dead speed figure and runs
a sheen over the bar. `stage`/`platform` are machine-readable; the wording
is localised in the frontend.

Finally, `modelscope.ai` is its own catalogue rather than an alias of
`modelscope.cn` — `referall13/EM1` exists only on `.ai` and
`jj3550945163/Krea-2-LORA` only on `.cn` — so its URLs were rejected with
"Invalid model URL format". Register it as `ModelScopeIntlSource`
(`platform="modelscope-ai"`, `msai:` group prefix, its own default
download directory) and derive every URL either deployment builds from a
per-class `base_url`. `modelscope.com` stays an alias of `.cn`, which is
what it redirects to. The frontend source table, the link dialog hints and
the docs mirror the split.

Verified against the live APIs: both reported `.ai` repositories list
their files, read their READMEs and yield name / version / base model /
trigger words / example images. Backend 3092 passed; frontend 1259 JS +
91 Vue passed. The nine locales carry the new progress copy in the next
commit.
This commit is contained in:
Will Miao
2026-09-17 07:47:07 +08:00
parent 1b1a8d63db
commit d572292142
34 changed files with 2385 additions and 135 deletions
+58 -29
View File
@@ -48,6 +48,7 @@ class PostProcessor:
readme_content: str = "",
source_context: Optional["ModelCardContext"] = None,
resolved_base_model: str = "",
metadata_source: str = "agent:enrich_hf_metadata",
) -> Dict[str, Any]:
"""Route *llm_output* to the correct skill post-processor.
@@ -63,13 +64,18 @@ class PostProcessor:
hints resolve to, used when the LLM did not supply one (which is the
normal case when the LLM was skipped).
*metadata_source* records who produced the metadata. The AI skill
keeps its historical value; the deterministic download-time hydration
passes its own so the two remain distinguishable. ``llm_enriched_at``
is only stamped when *llm_output* actually carries a provider answer.
Returns a dict with keys ``success`` (bool), ``updated_fields`` (list),
``preview_downloaded`` (bool), and ``errors`` (list).
"""
if skill_name == "enrich_hf_metadata":
return await self._process_enrich_hf_metadata(
model_path, llm_output, metadata, readme_content, source_context,
resolved_base_model,
resolved_base_model, metadata_source,
)
return {
"success": False,
@@ -89,6 +95,7 @@ class PostProcessor:
readme_content: str = "",
source_context: Optional["ModelCardContext"] = None,
resolved_base_model: str = "",
metadata_source: str = "agent:enrich_hf_metadata",
) -> Dict[str, Any]:
from ...metadata_ops import (
apply_metadata_updates,
@@ -135,6 +142,17 @@ class PostProcessor:
if new_base and self._should_overwrite(current_base, is_source_model):
updates["base_model"] = new_base
# model_name — the site's own display name, so a source download never
# shows up under its local filename. Written only while the name is
# still the untouched file stem: once a user renames a model that
# choice is theirs to keep.
site_name = ((source_context.model_name if source_context else "") or "").strip()
if is_source_model and site_name:
current_name = (metadata.get("model_name") or "").strip()
file_stem = (metadata.get("file_name") or "").strip()
if not current_name or current_name == file_stem:
updates["model_name"] = site_name
# trigger words → civitai.trainedWords
new_triggers = llm_output.get("trigger_words", [])
trigger_words_empty = True
@@ -142,14 +160,9 @@ class PostProcessor:
cleaned = [t.strip() for t in new_triggers if t.strip()]
cleaned = [t for t in cleaned if t.lower() not in ("none", "null", "n/a")]
trigger_words_empty = not cleaned
current_civitai = metadata.get("civitai") or {}
current_triggers = current_civitai.get("trainedWords") or []
current_triggers = (metadata.get("civitai") or {}).get("trainedWords") or []
if self._should_overwrite_list(current_triggers, is_source_model):
trig_civitai = dict(current_civitai)
if "civitai" in updates and isinstance(updates["civitai"], dict):
trig_civitai.update(updates["civitai"])
trig_civitai["trainedWords"] = cleaned
updates["civitai"] = trig_civitai
self._merge_civitai(updates, metadata, trainedWords=cleaned)
# modelDescription — the author's own summary (when the site keeps one
# outside the README, e.g. ModelScope's ``Description``) followed by the
@@ -175,12 +188,16 @@ class PostProcessor:
if not short_desc:
short_desc = site_description
if short_desc and is_source_model:
current_civitai = metadata.get("civitai") or {}
desc_civitai = dict(current_civitai)
if "civitai" in updates and isinstance(updates["civitai"], dict):
desc_civitai.update(updates["civitai"])
desc_civitai["description"] = short_desc
updates["civitai"] = desc_civitai
self._merge_civitai(updates, metadata, description=short_desc)
# The version label completes the card the way a CivitAI download does:
# the UI renders `civitai.name` as the version chip. It is per file,
# so a collection repository shows that checkpoint's own label.
site_version = (
(source_context.version_name if source_context else "") or ""
).strip()
if is_source_model and site_version:
self._merge_civitai(updates, metadata, name=site_version)
# gallery images → civitai.images (site example images, YAML frontmatter
# widget entries, and Sample Gallery markdown tables in the README body)
@@ -244,12 +261,7 @@ class PostProcessor:
all_images = _dedupe_images(site_images + readme_images)
if all_images:
gallery_images = all_images
current_civitai = metadata.get("civitai") or {}
gallery_civitai = dict(current_civitai)
if "civitai" in updates and isinstance(updates["civitai"], dict):
gallery_civitai.update(updates["civitai"])
gallery_civitai["images"] = all_images
updates["civitai"] = gallery_civitai
self._merge_civitai(updates, metadata, images=all_images)
# tags — the site's curated tags are authoritative content vocabulary, so
# they are kept alongside whatever the LLM proposed (the LLM is skipped
@@ -269,9 +281,12 @@ class PostProcessor:
if len(merged) > len(existing_tags) or is_source_model:
updates["tags"] = merged
# metadata_source & llm_enriched_at (always set)
updates["metadata_source"] = "agent:enrich_hf_metadata"
updates["llm_enriched_at"] = datetime.now(timezone.utc).isoformat()
# metadata_source is recorded for provenance; llm_enriched_at only means
# something when a provider actually answered, so the deterministic
# download-time hydration does not claim an enrichment that never ran.
updates["metadata_source"] = metadata_source
if llm_output:
updates["llm_enriched_at"] = datetime.now(timezone.utc).isoformat()
# LLM confidence, stored for the enrichment evaluation harness. The key
# must NOT start with an underscore: `BaseModelMetadata.from_dict()`
@@ -292,12 +307,7 @@ class PostProcessor:
if instance_prompt:
site_triggers = [instance_prompt]
if site_triggers:
current_civitai = metadata.get("civitai") or {}
trig_civitai = dict(current_civitai)
if "civitai" in updates and isinstance(updates["civitai"], dict):
trig_civitai.update(updates["civitai"])
trig_civitai["trainedWords"] = site_triggers
updates["civitai"] = trig_civitai
self._merge_civitai(updates, metadata, trainedWords=site_triggers)
preview_remote_url = (llm_output.get("preview_url") or "").strip()
# Fallback: if the LLM couldn't find a preview image in the cleaned
@@ -371,6 +381,25 @@ class PostProcessor:
"", "unknown",
)
@staticmethod
def _merge_civitai(
updates: Dict[str, Any], metadata: Dict[str, Any], **fields: Any
) -> None:
"""Layer *fields* onto the ``civitai`` block being assembled.
Description, version label, trigger words and gallery images all live
in the same dict and are contributed by separate branches, so each one
starts from what is already on disk and then applies whatever an
earlier branch queued in *updates*.
"""
merged = dict(metadata.get("civitai") or {})
queued = updates.get("civitai")
if isinstance(queued, dict):
merged.update(queued)
merged.update(fields)
updates["civitai"] = merged
@staticmethod
def _should_overwrite_list(current_list: List[str], is_source_model: bool) -> bool:
"""Return ``True`` when a list field should be overwritten."""