feat(download): fill model metadata from the source API on download

A ModelScope or Hugging Face download landed as a bare filename, hash and
source link; the model card stayed empty until the user ran "Enrich
Metadata with AI" by hand. But everything that makes a CivitAI download
useful — the display name, the description, the tags, the trigger words,
the example images, the preview — is already published by those sites'
public APIs, so asking for it at download time is deterministic work, not
model work.

Add `py/services/model_sources/hydration.py`, called by
`_save_source_metadata()` once the sidecar exists and the file is in the
scanner cache. It fetches the model card plus the site's card extras and
hands them to the same `PostProcessor` the AI skill uses, with an empty
`llm_output`, so the two paths cannot drift apart. What lands:

* `model_name` from the site's own display name (ModelScope's `Name`), so
  the card stops showing the local filename — written only while the value
  still equals the file stem, since once a user renames a model that
  choice is theirs to keep
* `civitai.name` from the matched version's label (`showName`), which the
  card renders as the version chip
* `civitai.description` / `modelDescription` from the author summary plus
  the README as HTML
* `civitai.images` / `preview_url` from the per-file example images
* `civitai.trainedWords` from the per-file trigger words
* `base_model`, `tags` and `usage_tips` as before

Provenance stays honest: the pass records
`metadata_source = "source:<platform>"` rather than the skill's
`agent:enrich_hf_metadata`, and — because no provider ran — it no longer
stamps `llm_enriched_at`; that stamp is now conditional on the LLM
actually answering, which is what the field means. The five hand-rolled
`civitai` dict merges in the post-processor collapse into one
`_merge_civitai()` helper.

Two guards keep it safe. Only a model whose stored
`source_platform`/`source_url` match the repository being downloaded is
updated, so a local file that merely shares a name never receives another
model's card; and a file already on disk is topped up too, which
back-fills models downloaded before this existed. READMEs and detail
payloads describe the repository rather than the file, so a short-lived
process-wide `ModelSourceCache` (300 s, 32 entries) keeps a batch over one
repository to two HTTP requests. Every failure is logged and swallowed:
hydration can never fail a download.

Fix the hash policy while here. `_save_source_metadata()` went straight to
`MetadataManager.create_default_metadata()`, bypassing the per-type
factory on the owning scanner, so a checkpoint paid a full SHA256 inside
the download request — `CheckpointScanner`/`OtherScanner` deliberately
record `hash_status="pending"` with an empty `sha256` for their multi-GB
files. Metadata is now created through `scanner._create_default_metadata()`.
Hydration copes with the empty hash: `_matching_versions()` falls back to
the repository basename, which is exactly what the download just wrote.

Report both post-transfer stages, which advance no byte counter and so
read as a stall: the bar sat at 100% showing `0 B/s` for the seconds spent
hashing and fetching. `_report_phase()` broadcasts
`{"status": "metadata", "stage": "indexing" | "source", "platform": ...}`,
and `LoadingManager` names the stage in the status line (keeping the batch
position), retitles the item line, replaces the dead speed figure and runs
a sheen over the bar. `stage`/`platform` are machine-readable; the wording
is localised in the frontend.

Finally, `modelscope.ai` is its own catalogue rather than an alias of
`modelscope.cn` — `referall13/EM1` exists only on `.ai` and
`jj3550945163/Krea-2-LORA` only on `.cn` — so its URLs were rejected with
"Invalid model URL format". Register it as `ModelScopeIntlSource`
(`platform="modelscope-ai"`, `msai:` group prefix, its own default
download directory) and derive every URL either deployment builds from a
per-class `base_url`. `modelscope.com` stays an alias of `.cn`, which is
what it redirects to. The frontend source table, the link dialog hints and
the docs mirror the split.

Verified against the live APIs: both reported `.ai` repositories list
their files, read their READMEs and yield name / version / base model /
trigger words / example images. Backend 3092 passed; frontend 1259 JS +
91 Vue passed. The nine locales carry the new progress copy in the next
commit.
This commit is contained in:
Will Miao
2026-09-17 07:47:07 +08:00
parent 1b1a8d63db
commit d572292142
34 changed files with 2385 additions and 135 deletions
+10 -1
View File
@@ -24,7 +24,12 @@ from .base import (
is_valid_source_id,
)
from .huggingface import HuggingFaceSource
from .modelscope import ModelScopeSource
from .hydration import (
hydrate_from_source,
load_model_card,
resolve_site_base_model,
)
from .modelscope import ModelScopeIntlSource, ModelScopeSource
from .registry import (
LEGACY_HF_URL_FIELD,
SOURCE_PLATFORM_FIELD,
@@ -52,6 +57,7 @@ __all__ = [
"ModelSourceCache",
"ModelSourceError",
"HuggingFaceSource",
"ModelScopeIntlSource",
"ModelScopeSource",
"SOURCE_PLATFORM_FIELD",
"SOURCE_URL_FIELD",
@@ -68,9 +74,12 @@ __all__ = [
"get_source",
"get_source_platform",
"has_external_source",
"hydrate_from_source",
"is_valid_source_id",
"list_sources",
"load_model_card",
"normalize_metadata_source",
"resolve_site_base_model",
"resolve_source_ref",
"source_group_key",
"source_label",
+30
View File
@@ -45,6 +45,7 @@ USER_AGENT = "ComfyUI-LoRA-Manager/1.0"
GROUP_PREFIXES: dict[str, str] = {
"huggingface": "hf",
"modelscope": "ms",
"modelscope-ai": "msai",
"tensorart": "ta",
}
@@ -77,6 +78,30 @@ class ModelCardContext:
description: str = ""
"""Author-written summary shown on the model page, outside the README."""
model_name: str = ""
"""Site-published display name for the repository.
Sites publish this next to the repository id (ModelScope's ``Name``).
It is what a CivitAI download would store as the model's name, so the
card never has to fall back to the local filename.
"""
model_name_localized: str = ""
"""Site-published localized name (ModelScope's ``ChineseName``)."""
version_name: str = ""
"""Site-published label for the requested file's version.
Resolved per file, like :attr:`example_images`: a repository publishes
one label per checkpoint (ModelScope's ``modelVersion.showName``).
"""
license: str = ""
"""License the site records for the repository."""
model_type: str = ""
"""Site-reported model type, e.g. ModelScope's ``AigcType`` (``LoRA``)."""
base_model: str = ""
"""Base model as reported by the site (possibly a site-local id)."""
@@ -104,6 +129,11 @@ class ModelCardContext:
return not any(
(
self.description,
self.model_name,
self.model_name_localized,
self.version_name,
self.license,
self.model_type,
self.base_model,
self.base_model_aliases,
self.official_tags,
+235
View File
@@ -0,0 +1,235 @@
"""Deterministic metadata hydration for freshly downloaded source models.
A CivitAI download writes a fully-populated metadata sidecar as part of the
download itself: the name, the description, the tags, the trigger words and
the example images all arrive with the file. A download from an external
model source (ModelScope, Hugging Face) has the same information behind a
public API, but historically landed as a bare filename plus a source URL that
the user had to enrich by hand ("Enrich Metadata with AI").
This module closes that gap without involving an LLM. It fetches the linked
site's model card, hands it to the same :class:`~py.services.agent.post_processor.PostProcessor`
the AI skill uses, and writes the result. Everything it applies is data the
site published, so it is safe to run automatically on every download and to
treat as a fallback for the gaps the LLM would otherwise fill.
Nothing here may break a download: every failure is logged and normalised to
"the site had nothing to contribute".
"""
from __future__ import annotations
import logging
import os
import time
from typing import TYPE_CHECKING, Optional
from .base import ModelCardContext, ModelSourceCache
from .registry import get_source, resolve_source_ref
if TYPE_CHECKING: # pragma: no cover - typing only
from .base import ModelSource, SourceRef
logger = logging.getLogger(__name__)
#: How long a fetched repository payload stays usable. A download batch walks
#: a repository's files one HTTP request at a time, and the README plus the
#: detail payload describe the *repository*, not the file, so re-fetching them
#: per file would be pure waste. They expire so an edited model card is still
#: picked up by the next batch.
SHARED_CACHE_TTL = 300.0
#: Upper bound on memoised repositories; a long-running server must not grow
#: without limit.
SHARED_CACHE_MAX_ENTRIES = 32
#: ``"<platform>:<source_id>"`` → ``(expiry, memo)``.
_shared_caches: dict[str, tuple[float, ModelSourceCache]] = {}
def shared_source_cache(platform: str, source_id: str) -> ModelSourceCache:
"""Return a short-lived per-repository memo for download-time hydration."""
now = time.monotonic()
key = f"{platform}:{source_id}"
entry = _shared_caches.get(key)
if entry is not None and entry[0] > now:
return entry[1]
for expired in [k for k, (expiry, _) in _shared_caches.items() if expiry <= now]:
_shared_caches.pop(expired, None)
if len(_shared_caches) >= SHARED_CACHE_MAX_ENTRIES:
oldest = min(_shared_caches, key=lambda k: _shared_caches[k][0])
_shared_caches.pop(oldest, None)
cache = ModelSourceCache()
_shared_caches[key] = (now + SHARED_CACHE_TTL, cache)
return cache
def reset_shared_caches() -> None:
"""Drop every memoised repository — used by tests."""
_shared_caches.clear()
async def load_model_card(
source: "ModelSource",
source_id: str,
cache: Optional[ModelSourceCache] = None,
) -> str:
"""Return *source_id*'s README, reusing *cache* when one is supplied.
Only successful reads are memoised, leaving a transient failure to be
retried for the next file of the same repository.
"""
key = f"{source.platform}:{source_id}"
if cache is not None:
cached = cache.readmes.get(key)
if cached is not None:
return cached
readme = await source.fetch_model_card(source_id)
if cache is not None and readme:
cache.readmes[key] = readme
return readme or ""
async def resolve_site_base_model(context: ModelCardContext) -> str:
"""Resolve the site's base-model hints to a canonical name, or ``""``.
Sites name base models in their own vocabulary (ModelScope publishes both
``krea/Krea-2-Turbo`` and the ``KREA_2_TURBO`` enum). The resolver is
strict and only ever returns a name the canonical vocabulary already
contains, so an uncertain hint yields ``""`` rather than a plausible-looking
wrong value.
"""
hints = [*context.base_model_aliases, context.base_model]
if not any(hints):
return ""
# Imported lazily: pulling in the agent package at module scope would make
# the model-source package import itself while it is still initialising.
try:
from ...metadata_ops import list_base_models
from ..agent.base_model_resolver import resolve_base_model
known_names = await list_base_models()
except Exception as exc:
logger.warning("Could not resolve a site base model: %s", exc)
return ""
return resolve_base_model(hints, known_names)
async def hydrate_from_source(
file_path: str,
*,
ref: "SourceRef",
cache: Optional[ModelSourceCache] = None,
) -> list[str]:
"""Apply the linked site's published metadata to a downloaded model.
This is the deterministic counterpart of the ``enrich_hf_metadata`` skill:
it produces the same populated model card a CivitAI download produces,
without an LLM and without user action.
Args:
file_path: The just-downloaded model file, whose sidecar already
carries the SHA256 used to match the right file in a collection
repository.
ref: The source the file came from.
cache: Optional per-call memo; defaults to a short-lived shared one so
a batch over one repository fetches its card only once.
Returns:
The names of the metadata fields that changed. Never raises — a site
that is down, or an API that changed shape, must not fail a download.
"""
try:
source = get_source(ref.platform)
if source is None or not source.supports_enrichment:
return []
from ...metadata_ops import read_metadata
metadata = await read_metadata(file_path)
if not metadata:
logger.debug("No metadata to hydrate for %s", file_path)
return []
# Only a model that is actually linked to this repository may be
# updated. The download path writes those fields just before calling
# us; a file that merely shares a name with the requested one must not
# be given another model's card.
linked = resolve_source_ref(metadata)
if linked is None or (linked.platform, linked.source_id) != (
ref.platform,
ref.source_id,
):
logger.debug(
"Not hydrating %s: linked to %s, not %s",
file_path, linked.url if linked else "no model source", ref.url,
)
return []
memo = cache if cache is not None else shared_source_cache(
ref.platform, ref.source_id
)
readme = await load_model_card(source, ref.source_id, memo)
context = await source.fetch_model_card_context(
ref.source_id,
os.path.basename(file_path),
sha256=(metadata.get("sha256") or "").strip(),
cache=memo,
)
if context.is_empty() and not readme:
logger.debug(
"No published metadata for %s on %s", ref.source_id, ref.platform
)
return []
resolved_base_model = await resolve_site_base_model(context)
from ..agent.post_processor import PostProcessor
result = await PostProcessor().process(
skill_name="enrich_hf_metadata",
model_path=file_path,
llm_output={},
metadata=metadata,
readme_content=readme,
source_context=context,
resolved_base_model=resolved_base_model,
metadata_source=f"source:{ref.platform}",
)
if not result.get("success", True):
logger.debug(
"Hydration reported failure for %s: %s",
file_path, result.get("errors"),
)
return []
updated = list(result.get("updated_fields") or [])
logger.info(
"Hydrated %s from %s (%s): %s",
file_path, source.label or ref.platform, ref.source_id,
", ".join(updated) or "nothing to change",
)
return updated
except Exception as exc: # pragma: no cover - defensive by design
logger.warning("Source hydration failed for %s: %s", file_path, exc)
return []
__all__ = [
"SHARED_CACHE_MAX_ENTRIES",
"SHARED_CACHE_TTL",
"hydrate_from_source",
"load_model_card",
"reset_shared_caches",
"resolve_site_base_model",
"shared_source_cache",
]
+169 -38
View File
@@ -1,4 +1,4 @@
"""ModelScope (魔搭社区) model source.
"""ModelScope (魔搭社区) model sources.
ModelScope exposes the same "model card as README.md" convention as
Hugging Face, including a YAML frontmatter block that often carries
@@ -10,11 +10,13 @@ none of which requires an API key for public models:
the same content through the API, used as a fallback when the resolve
URL is unavailable.
* ``/api/v1/models/{owner}/{name}`` — the model-detail payload behind the
model page. It carries the author's summary (``Description``), the
site-curated tags (``OfficialTags``), and, per published version, the
model filenames (``MuseInfo.versions[].stats.fileList``) together with
that file's example images (``coverImages``) and trigger words. See
:meth:`ModelScopeSource.fetch_model_card_context`.
model page. It carries the repository's display name (``Name`` /
``ChineseName``), the author's summary (``Description``), the license, the
AIGC type, the site tags (``OfficialTags``, falling back to ``Tags``), and,
per published version, the model filenames
(``MuseInfo.versions[].stats.fileList``) together with that version's label
(``modelVersion.showName``), example images (``coverImages``) and trigger
words. See :meth:`ModelScopeSource.fetch_model_card_context`.
* ``/api/v1/models/{owner}/{name}/repo/files?Revision=..`` — the file
listing backing the download picker. It reports real sizes for LFS
files (not the pointer size), so no extra HEAD request is needed.
@@ -28,6 +30,12 @@ valid; the CDN URL must never be cached.
The README and the detail payload both describe the whole repository rather
than one file, so a per-run ``ModelSourceCache`` keeps them from being read
again for every checkpoint of a collection repository.
Two deployments are served by this module. ``modelscope.cn`` (with
``modelscope.com`` as a redirect alias) and ``modelscope.ai`` are *separate
catalogues*, not mirrors, so they are registered as distinct sources:
:class:`ModelScopeSource` and :class:`ModelScopeIntlSource`. Every URL either
class builds is derived from its ``base_url``.
"""
from __future__ import annotations
@@ -36,7 +44,7 @@ import json
import logging
import os
import re
from typing import TYPE_CHECKING, Any, Optional
from typing import TYPE_CHECKING, Any, Iterable, Optional
from .base import (
ModelCardContext,
@@ -52,18 +60,28 @@ if TYPE_CHECKING: # pragma: no cover - typing only
logger = logging.getLogger(__name__)
_URL_PATTERN = re.compile(
r"https?://(?:www\.)?modelscope\.(?:cn|com)/models/(?P<id>[^/?#\s]+/[^/?#\s]+)"
)
#: ModelScope runs two independent catalogues. ``modelscope.com`` is a
#: redirect alias of the mainland site, but ``modelscope.ai`` is the
#: *international* deployment with its own repository catalogue — a repository
#: published on one is routinely absent from the other (``referall13/EM1``
#: exists only on ``.ai``, ``jj3550945163/Krea-2-LORA`` only on ``.cn``). The
#: host therefore decides which site, API and CDN a model belongs to, and the
#: two deployments are registered as separate sources rather than folded into
#: one id.
_MAINLAND_HOSTS = r"modelscope\.(?:cn|com)"
_INTERNATIONAL_HOSTS = r"modelscope\.ai"
#: Trailing view segments the site appends to a model URL; accepted verbatim
#: when the user pastes a browser tab URL.
_VIEW_SEGMENTS = r"(?:summary|files|model-file|readme|community|evaluation)?"
_STRICT_URL_PATTERN = re.compile(
r"https?://(?:www\.)?modelscope\.(?:cn|com)/models/(?P<id>[^/?#\s]+/[^/?#\s]+)"
rf"/?{_VIEW_SEGMENTS}/?$"
)
def _url_patterns(hosts: str) -> tuple[re.Pattern[str], re.Pattern[str]]:
"""Build the lenient and strict model-URL patterns for *hosts*."""
body = rf"https?://(?:www\.)?(?:{hosts})/models/(?P<id>[^/?#\s]+/[^/?#\s]+)"
return re.compile(body), re.compile(rf"{body}/?{_VIEW_SEGMENTS}/?$")
#: ``master`` is ModelScope's default branch; ``main`` is tried as a fallback
#: for repos imported from Hugging Face.
@@ -71,7 +89,12 @@ _REVISIONS = ("master", "main")
class ModelScopeSource(ModelSource):
"""ModelScope (``modelscope.cn``)."""
"""ModelScope's mainland site (``modelscope.cn``).
``modelscope.com`` is accepted as an alias of it. The international
deployment is :class:`ModelScopeIntlSource`; everything below is written in
terms of ``base_url`` so both share one implementation.
"""
platform = "modelscope"
label = "ModelScope"
@@ -79,15 +102,18 @@ class ModelScopeSource(ModelSource):
supports_download = True
default_revision = "master"
default_subdir = "modelscope"
url_pattern = _URL_PATTERN
strict_url_pattern = _STRICT_URL_PATTERN
#: Origin every outgoing URL is built from.
base_url = "https://modelscope.cn"
url_pattern, strict_url_pattern = _url_patterns(_MAINLAND_HOSTS)
def canonical_url(self, source_id: str) -> str:
return f"https://modelscope.cn/models/{source_id}"
return f"{self.base_url}/models/{source_id}"
def asset_base_url(self, source_id: str, revision: str = "") -> str:
return (
f"https://modelscope.cn/models/{source_id}/resolve/"
f"{self.base_url}/models/{source_id}/resolve/"
f"{self.resolve_revision(revision)}"
)
@@ -96,7 +122,7 @@ class ModelScopeSource(ModelSource):
for revision in _REVISIONS:
text = await fetch_text(
f"https://modelscope.cn/models/{source_id}/resolve/{revision}/README.md"
f"{self.base_url}/models/{source_id}/resolve/{revision}/README.md"
)
if text:
return text
@@ -105,7 +131,7 @@ class ModelScopeSource(ModelSource):
# environments where the CDN resolve host is blocked.
for revision in _REVISIONS:
text = await fetch_text(
"https://modelscope.cn/api/v1/models/"
f"{self.base_url}/api/v1/models/"
f"{source_id}/repo?Revision={revision}&FilePath=README.md"
)
if text:
@@ -158,7 +184,7 @@ class ModelScopeSource(ModelSource):
return cache.provider[cache_key]
status, payload = await fetch_json(
f"https://modelscope.cn/api/v1/models/{source_id}"
f"{self.base_url}/api/v1/models/{source_id}"
)
if status != 200 or not isinstance(payload, dict):
logger.debug(
@@ -185,7 +211,7 @@ class ModelScopeSource(ModelSource):
revision = self.resolve_revision(revision)
status, payload = await fetch_json(
"https://modelscope.cn/api/v1/models/"
f"{self.base_url}/api/v1/models/"
f"{source_id}/repo/files?Revision={revision}"
)
@@ -208,18 +234,37 @@ class ModelScopeSource(ModelSource):
self, source_id: str, filename: str, revision: str = ""
) -> str:
return (
f"https://modelscope.cn/models/{source_id}/resolve/"
f"{self.base_url}/models/{source_id}/resolve/"
f"{self.resolve_revision(revision)}/{filename}"
)
def page_url_for_file(self, source_id: str, filename: str) -> str:
return (
f"https://modelscope.cn/models/{source_id}/file/view/"
f"{self.base_url}/models/{source_id}/file/view/"
f"{self.default_revision}/{filename}"
)
__all__ = ["ModelScopeSource"]
class ModelScopeIntlSource(ModelScopeSource):
"""ModelScope's international site (``modelscope.ai``).
A separate catalogue rather than a mirror, so it is registered under its
own platform id: the two deployments must not share a version group, a
"use default paths" directory, or a stored ``source_url``. The detail API,
the file listing, the resolve URLs and the CDN redirect all behave exactly
like the mainland site, which is why every URL here is derived from
:attr:`base_url` instead of being duplicated.
"""
platform = "modelscope-ai"
label = "ModelScope (International)"
default_subdir = "modelscope-ai"
base_url = "https://www.modelscope.ai"
url_pattern, strict_url_pattern = _url_patterns(_INTERNATIONAL_HOSTS)
__all__ = ["ModelScopeIntlSource", "ModelScopeSource"]
# ---------------------------------------------------------------------------
@@ -229,6 +274,33 @@ __all__ = ["ModelScopeSource"]
#: Trigger-word values that mean "the author left this blank".
_EMPTY_TRIGGER_VALUES = frozenset({"none", "null", "n/a"})
#: Repository tags that only restate what the model *is* (its library, task or
#: framework) rather than what it depicts. ModelScope mixes both into the
#: plain ``Tags`` list, and a card tagged "lora" or "text-to-image" is noise.
_GENERIC_TAGS = frozenset(
{
"any-to-any",
"checkpoint",
"controlnet",
"diffusers",
"embedding",
"image-text-to-text",
"image-to-image",
"image-to-video",
"lora",
"lycoris",
"onnx",
"pytorch",
"safetensors",
"tensorflow",
"text-to-image",
"text-to-speech",
"text-to-video",
"textual-inversion",
"vae",
}
)
def _clean_text(value: Any) -> str:
"""Return a stripped string for *value*, or ``""`` for anything else."""
@@ -259,9 +331,13 @@ def _build_card_context(
context = ModelCardContext(
description=_clean_text(data.get("Description")),
model_name=_clean_text(data.get("Name")),
model_name_localized=_clean_text(data.get("ChineseName")),
license=_clean_text(data.get("License")),
model_type=_clean_text(data.get("AigcType")),
base_model=_first_string(data.get("BaseModel")),
base_model_aliases=_base_model_aliases(data),
official_tags=_official_tags(data.get("OfficialTags")),
official_tags=_official_tags(data),
)
versions = _matching_versions(
@@ -271,6 +347,7 @@ def _build_card_context(
sha256=sha256,
)
if versions:
context.version_name = _version_label(versions)
context.example_images = _cover_image_urls(versions)
context.trigger_words = _version_trigger_words(versions)
return context
@@ -303,26 +380,63 @@ def _base_model_aliases(data: dict[str, Any]) -> list[str]:
return aliases
def _official_tags(value: Any) -> list[str]:
"""Extract the site-curated tag values from ``OfficialTags``.
def _official_tags(data: dict[str, Any]) -> list[str]:
"""Return the content tags the site publishes for the repository.
ModelScope's entries are dicts carrying an English ``Tag`` plus a
``ChineseName``; the English value is the curated content vocabulary, so
that is the one surfaced here.
``OfficialTags`` is ModelScope's curated content vocabulary and is
preferred whenever it is populated. Plenty of AIGC repositories leave it
empty and carry only the plain ``Tags`` list, which mixes content tags with
framework and task categories; those categories are dropped so a card is
not handed "lora" and "text-to-image" as if they described the model.
"""
curated = _dedupe(_tag_values(data.get("OfficialTags")))
if curated:
return curated
generic = set(_GENERIC_TAGS)
for value in (
data.get("AigcType"),
data.get("Libraries"),
data.get("Frameworks"),
):
for item in value if isinstance(value, list) else [value]:
text = _clean_text(item).lower()
if text:
generic.add(text)
return _dedupe(
tag for tag in _tag_values(data.get("Tags")) if tag.lower() not in generic
)
def _tag_values(value: Any) -> list[str]:
"""Return the tag strings from either shape ModelScope publishes.
``OfficialTags`` is a list of ``{"Tag": ..., "ChineseName": ...}`` dicts
carrying an English value; the plain ``Tags`` list is already strings.
"""
tags: list[str] = []
if not isinstance(value, list):
return tags
return []
tags: list[str] = []
for entry in value:
if not isinstance(entry, dict):
continue
tag = _clean_text(entry.get("Tag"))
if tag and tag not in tags:
tag = _clean_text(entry.get("Tag") if isinstance(entry, dict) else entry)
if tag:
tags.append(tag)
return tags
def _dedupe(values: Iterable[str]) -> list[str]:
"""Drop empties and repeats, keeping the first spelling seen."""
unique: list[str] = []
for value in values:
if value and value not in unique:
unique.append(value)
return unique
def _version_files(version: dict[str, Any]) -> list[str]:
"""Return the model filenames covered by one ``MuseInfo.versions`` entry.
@@ -359,6 +473,23 @@ def _version_show_name(version: dict[str, Any]) -> str:
return _clean_text(model_version.get("showName")).lower()
def _version_label(versions: list[dict[str, Any]]) -> str:
"""Return the first published version label, preserving its spelling.
Unlike :func:`_version_show_name` this is for display, so the label is
not lowercased.
"""
for version in versions:
model_version = version.get("modelVersion")
if not isinstance(model_version, dict):
continue
label = _clean_text(model_version.get("showName"))
if label:
return label
return ""
def _file_digests(data: dict[str, Any]) -> dict[str, str]:
"""Return ``basename -> sha256`` for every published weight file.
+4 -1
View File
@@ -13,15 +13,18 @@ from typing import Any, Dict, Mapping, Optional
from .base import GROUP_PREFIXES, ModelSource, SourceRef, clean_source_url
from .huggingface import HuggingFaceSource
from .modelscope import ModelScopeSource
from .modelscope import ModelScopeIntlSource, ModelScopeSource
from .tensorart import TensorArtSource
logger = logging.getLogger(__name__)
#: Order matters only for disambiguation; the URL patterns are disjoint.
#: ``modelscope.ai`` is a separate catalogue from ``modelscope.cn`` rather than
#: an alias, which is why it gets its own entry (see ``modelscope.py``).
_SOURCES: tuple[ModelSource, ...] = (
HuggingFaceSource(),
ModelScopeSource(),
ModelScopeIntlSource(),
TensorArtSource(),
)