Source profileQuality 68/100

affaan-m/ECC/docs/ja-JP/skills/content-hash-cache-pattern/SKILL.md

content-hash-cache-pattern

Review content-hash-cache-pattern's use cases, installation, workflow, and original source instructions.

Source repository stars
234,327
Declared platforms
0
Static risk flags
0
Last source update
2026-07-27
Source checked
2026-07-28

Decision brief

What it does—and where it fits

SHA-256コンテンツハッシュをキャッシュキーとして使用して、高コストなファイル処理結果(PDF解析、テキスト抽出、画像分析)をキャッシュします。パスベースのキャッシュとは異なり、このアプローチはファイルの移動/名前変更に対して生き残り、コンテンツが変更されたときに自動的に無効化されます。

Best for

    Not for

    • Tasks that require unconfirmed production actions or broad system permissions.
    • Environments where the pinned source and install steps cannot be inspected.

    Compatibility matrix

    Platform support, with evidence labels

    PlatformStatusEvidenceWhat to check
    CodexNot declaredNo explicit evidencePortability before use
    Claude CodeNot declaredNo explicit evidencePortability before use
    CursorNot declaredNo explicit evidencePortability before use
    Gemini CLINot declaredNo explicit evidencePortability before use
    Open the compatibility checker

    Installation

    Inspect first. Install second.

    The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

    Source-detected install commandSource
    npx skills add https://github.com/affaan-m/ECC --skill "docs/ja-JP/skills/content-hash-cache-pattern"
    Safe inspection promptEditorial

    Inspect the Agent Skill "content-hash-cache-pattern" from https://github.com/affaan-m/ECC/blob/4e973d3eaf92d97f8d2e2d8abb39d8bdc8711b38/docs/ja-JP/skills/content-hash-cache-pattern/SKILL.md at commit 4e973d3eaf92d97f8d2e2d8abb39d8bdc8711b38. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

    Workflow

    What the source asks the agent to do

    1. 01

      起動条件

      ファイル処理パイプラインの構築(PDF、画像、テキスト抽出)

      ファイル処理パイプラインの構築(PDF、画像、テキスト抽出)処理コストが高く、同じファイルが繰り返し処理される場合--cache/--no-cacheCLIオプションが必要な場合
    2. 02

      コアパターン

      パスではなくファイルコンテンツをキャッシュキーとして使用します:

      パスではなくファイルコンテンツをキャッシュキーとして使用します:なぜコンテンツハッシュ? ファイルの名前変更/移動 = キャッシュヒット。コンテンツ変更 = 自動無効化。インデックスファイル不要。各キャッシュエントリは{hash}.jsonとして保存されます — ハッシュによるO(1)検索、インデックスファイル不要。
    3. 03

      1. コンテンツハッシュベースのキャッシュキー

      パスではなくファイルコンテンツをキャッシュキーとして使用します:

      パスではなくファイルコンテンツをキャッシュキーとして使用します:なぜコンテンツハッシュ? ファイルの名前変更/移動 = キャッシュヒット。コンテンツ変更 = 自動無効化。インデックスファイル不要。
    4. 04

      2. キャッシュエントリの凍結データクラス

      Review the “2. キャッシュエントリの凍結データクラス” section in the pinned source before continuing.

      Review and apply the “2. キャッシュエントリの凍結データクラス” source section.

    Permission review

    Static risk signals and limitations

    No configured static risk pattern was detected

    This is not proof of safety. Runtime behavior, indirect dependencies, and hidden external systems are outside the static scan.

    Evidence record

    Why each signal appears

    EvidenceSourceComputedTestedEditorial
    SignalValueEvidence typeMeaning
    Quality score68/100ComputedDocumentation, specificity, maintenance, and trust rules
    Repository stars234,327SourceRepository attention, not individual Skill quality
    Compatibility0 platformsSourceDeclared in the catalog source record
    Usage guideautomated source guideEditorialGenerated or reviewed according to the visible evidence level

    Pinned source

    Provenance and original SKILL.md

    Repository
    affaan-m/ECC
    Skill path
    docs/ja-JP/skills/content-hash-cache-pattern/SKILL.md
    Commit
    4e973d3eaf92d97f8d2e2d8abb39d8bdc8711b38
    License
    MIT
    Collected
    2026-07-28
    Default branch
    main
    View the original SKILL.md

    コンテンツハッシュファイルキャッシュパターン

    SHA-256コンテンツハッシュをキャッシュキーとして使用して、高コストなファイル処理結果(PDF解析、テキスト抽出、画像分析)をキャッシュします。パスベースのキャッシュとは異なり、このアプローチはファイルの移動/名前変更に対して生き残り、コンテンツが変更されたときに自動的に無効化されます。

    起動条件

    • ファイル処理パイプラインの構築(PDF、画像、テキスト抽出)
    • 処理コストが高く、同じファイルが繰り返し処理される場合
    • --cache/--no-cacheCLIオプションが必要な場合
    • 既存の純粋な関数を変更せずにキャッシュを追加したい場合

    コアパターン

    1. コンテンツハッシュベースのキャッシュキー

    パスではなくファイルコンテンツをキャッシュキーとして使用します:

    import hashlib
    from pathlib import Path
    
    _HASH_CHUNK_SIZE = 65536  # 大きなファイルには64KBチャンク
    
    def compute_file_hash(path: Path) -> str:
        """ファイルコンテンツのSHA-256(大きなファイルにはチャンク処理)。"""
        if not path.is_file():
            raise FileNotFoundError(f"File not found: {path}")
        sha256 = hashlib.sha256()
        with open(path, "rb") as f:
            while True:
                chunk = f.read(_HASH_CHUNK_SIZE)
                if not chunk:
                    break
                sha256.update(chunk)
        return sha256.hexdigest()
    

    なぜコンテンツハッシュ? ファイルの名前変更/移動 = キャッシュヒット。コンテンツ変更 = 自動無効化。インデックスファイル不要。

    2. キャッシュエントリの凍結データクラス

    from dataclasses import dataclass
    
    @dataclass(frozen=True, slots=True)
    class CacheEntry:
        file_hash: str
        source_path: str
        document: ExtractedDocument  # キャッシュされた結果
    

    3. ファイルベースのキャッシュストレージ

    各キャッシュエントリは{hash}.jsonとして保存されます — ハッシュによるO(1)検索、インデックスファイル不要。

    import json
    from typing import Any
    
    def write_cache(cache_dir: Path, entry: CacheEntry) -> None:
        cache_dir.mkdir(parents=True, exist_ok=True)
        cache_file = cache_dir / f"{entry.file_hash}.json"
        data = serialize_entry(entry)
        cache_file.write_text(json.dumps(data, ensure_ascii=False), encoding="utf-8")
    
    def read_cache(cache_dir: Path, file_hash: str) -> CacheEntry | None:
        cache_file = cache_dir / f"{file_hash}.json"
        if not cache_file.is_file():
            return None
        try:
            raw = cache_file.read_text(encoding="utf-8")
            data = json.loads(raw)
            return deserialize_entry(data)
        except (json.JSONDecodeError, ValueError, KeyError):
            return None  # 破損をキャッシュミスとして扱う
    

    4. サービスレイヤーラッパー(SRP)

    処理関数を純粋に保ちます。キャッシュを別のサービスレイヤーとして追加します。

    def extract_with_cache(
        file_path: Path,
        *,
        cache_enabled: bool = True,
        cache_dir: Path = Path(".cache"),
    ) -> ExtractedDocument:
        """サービスレイヤー: キャッシュチェック -> 抽出 -> キャッシュ書き込み。"""
        if not cache_enabled:
            return extract_text(file_path)  # 純粋な関数、キャッシュの知識なし
    
        file_hash = compute_file_hash(file_path)
    
        # キャッシュを確認
        cached = read_cache(cache_dir, file_hash)
        if cached is not None:
            logger.info("Cache hit: %s (hash=%s)", file_path.name, file_hash[:12])
            return cached.document
    
        # キャッシュミス -> 抽出 -> 保存
        logger.info("Cache miss: %s (hash=%s)", file_path.name, file_hash[:12])
        doc = extract_text(file_path)
        entry = CacheEntry(file_hash=file_hash, source_path=str(file_path), document=doc)
        write_cache(cache_dir, entry)
        return doc
    

    主要な設計上の決定

    決定根拠
    SHA-256コンテンツハッシュパス非依存、コンテンツ変更で自動無効化
    {hash}.jsonファイル命名O(1)検索、インデックスファイル不要
    サービスレイヤーラッパーSRP: 抽出は純粋に保ち、キャッシュは別の関心事
    手動JSONシリアル化凍結データクラスのシリアル化を完全制御
    破損はNoneを返すグレースフルデグラデーション、次回の実行で再処理
    cache_dir.mkdir(parents=True)最初の書き込み時に遅延ディレクトリ作成

    ベストプラクティス

    • パスではなくコンテンツをハッシュ — パスは変わるが、コンテンツのアイデンティティは変わらない
    • 大きなファイルはチャンク処理でハッシュ — ファイル全体をメモリに読み込まないようにする
    • 処理関数を純粋に保つ — キャッシュについて何も知らないようにする
    • 切り捨てたハッシュでキャッシュヒット/ミスをログ記録 — デバッグのため
    • 破損をグレースフルに処理 — 無効なキャッシュエントリはミスとして扱い、クラッシュしない

    避けるべきアンチパターン

    # 悪い例: パスベースのキャッシュ(ファイルの移動/名前変更で壊れる)
    cache = {"/path/to/file.pdf": result}
    
    # 悪い例: 処理関数内にキャッシュロジックを追加(SRP違反)
    def extract_text(path, *, cache_enabled=False, cache_dir=None):
        if cache_enabled:  # この関数は今や2つの責任を持っている
            ...
    
    # 悪い例: ネストされた凍結データクラスでdataclasses.asdict()を使用
    # (複雑なネストされた型で問題を引き起こす可能性がある)
    data = dataclasses.asdict(entry)  # 代わりに手動シリアル化を使用
    

    使用すべき場合

    • ファイル処理パイプライン(PDF解析、OCR、テキスト抽出、画像分析)
    • --cache/--no-cacheオプションが有益なCLIツール
    • 同じファイルが複数回にわたって現れるバッチ処理
    • 既存の純粋な関数を変更せずにキャッシュを追加する場合

    使用すべきでない場合

    • 常に最新でなければならないデータ(リアルタイムフィード)
    • 非常に大きなキャッシュエントリ(代わりにストリーミングを検討)
    • ファイルコンテンツ以外のパラメータに依存する結果(例:異なる抽出設定)

    Alternatives

    Compare before choosing