Source profileQuality 80/100

xuzhougeng/wisp-science/skills/openfold3/SKILL.md

openfold3

Structure prediction using OpenFold3, an open-weights PyTorch reproduction of AlphaFold3 from the AlQuraishi Lab. Use this skill when predicting protein/nucleic-acid/ligand complex structures with an Apache-2.0-licensed AF3 reimplementation.

Source repository stars
560
Declared platforms
0
Static risk flags
1
Last source update
2026-07-28
Source checked
2026-07-28

Decision brief

What it does—and where it fits

Structure prediction using OpenFold3, an open-weights PyTorch reproduction of AlphaFold3 from the AlQuraishi Lab. 0-licensed AF3 reimplementation.

Best for

  • Use this skill when predicting protein/nucleic-acid/ligand complex structures with an Apache-2.

Not for

  • Tasks that require unconfirmed production actions or broad system permissions.
  • Environments where the pinned source and install steps cannot be inspected.

Compatibility matrix

Platform support, with evidence labels

PlatformStatusEvidenceWhat to check
CodexNot declaredNo explicit evidencePortability before use
Claude CodeNot declaredNo explicit evidencePortability before use
CursorNot declaredNo explicit evidencePortability before use
Gemini CLINot declaredNo explicit evidencePortability before use
Open the compatibility checker

Installation

Inspect first. Install second.

The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

Source-detected install commandSource
npx skills add https://github.com/xuzhougeng/wisp-science --skill "skills/openfold3"
Safe inspection promptEditorial

Inspect the Agent Skill "openfold3" from https://github.com/xuzhougeng/wisp-science/blob/95d2c13d1665d46a388b5bdc998dcce0d5ec2eee/skills/openfold3/SKILL.md at commit 95d2c13d1665d46a388b5bdc998dcce0d5ec2eee. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

Workflow

What the source asks the agent to do

  1. 01

    How to run

    The default attention kernel is DeepSpeed DS4SciEvoformerAttention. If DeepSpeed is unavailable, switch to the cuEquivariance triangle kernels (no build-from-source) by overriding the eval memory settings in modelconfig.py (usedeepspeedevoattention: False, usecueqtrianglekernels…

    The default attention kernel is DeepSpeed DS4SciEvoformerAttention. If DeepSpeed is unavailable, switch to the cuEquivariance triangle kernels (no build-from-source) by overriding the eval memory settings in modelconfig…Apache-2.0, 2.3 GB from HF OpenFold/OpenFold3. The repo is gated (auto-approval) — accept the access form on the HF model page and authenticate (huggingface-cli login or HFTOKEN) before downloading:runopenfold will also auto-download to $OPENFOLDCACHE on first run if egress is open and HF credentials are available (either HFTOKEN or a prior huggingface-cli login) with repo access granted. The interactive setupopen…
  2. 02

    Prerequisites

    Review the “Prerequisites” section in the pinned source before continuing.

    Review and apply the “Prerequisites” source section.
  3. 03

    Installation

    The default attention kernel is DeepSpeed DS4SciEvoformerAttention. If DeepSpeed is unavailable, switch to the cuEquivariance triangle kernels (no build-from-source) by overriding the eval memory settings in modelconfig.py (usedeepspeedevoattention: False, usecueqtrianglekernels…

    The default attention kernel is DeepSpeed DS4SciEvoformerAttention. If DeepSpeed is unavailable, switch to the cuEquivariance triangle kernels (no build-from-source) by overriding the eval memory settings in modelconfig…
  4. 04

    Weights

    Apache-2.0, 2.3 GB from HF OpenFold/OpenFold3. The repo is gated (auto-approval) — accept the access form on the HF model page and authenticate (huggingface-cli login or HFTOKEN) before downloading:

    Apache-2.0, 2.3 GB from HF OpenFold/OpenFold3. The repo is gated (auto-approval) — accept the access form on the HF model page and authenticate (huggingface-cli login or HFTOKEN) before downloading:runopenfold will also auto-download to $OPENFOLDCACHE on first run if egress is open and HF credentials are available (either HFTOKEN or a prior huggingface-cli login) with repo access granted. The interactive setupopen…
  5. 05

    Running

    runopenfold discovers the checkpoint under $OPENFOLDCACHE automatically. Only pass --inference-ckpt-path if you have a non-standard layout or multiple checkpoints and need to pin one explicitly.

    runopenfold discovers the checkpoint under $OPENFOLDCACHE automatically. Only pass --inference-ckpt-path if you have a non-standard layout or multiple checkpoints and need to pin one explicitly.For MSA + templates (slower, higher accuracy), drop the two false flags. The MSA server is api.colabfold.com; template chain-ID remap hits data.rcsb.org (GraphQL) — both must be reachable.

Permission review

Static risk signals and limitations

Reads files

low · line 162

The documentation asks the agent to read local files, directories, or repositories.

| `libXrender.so.1: cannot open shared object file` | rdkit (via pdbeccdutils) needs X11 render libs | `apt-get install libxrender1 libxext6 libsm6` |

Evidence record

Why each signal appears

EvidenceSourceComputedTestedEditorial
SignalValueEvidence typeMeaning
Quality score80/100ComputedDocumentation, specificity, maintenance, and trust rules
Repository stars560SourceRepository attention, not individual Skill quality
Compatibility0 platformsSourceDeclared in the catalog source record
Usage guideautomated source guideEditorialGenerated or reviewed according to the visible evidence level

Pinned source

Provenance and original SKILL.md

Repository
xuzhougeng/wisp-science
Skill path
skills/openfold3/SKILL.md
Commit
95d2c13d1665d46a388b5bdc998dcce0d5ec2eee
License
AGPL-3.0
Collected
2026-07-28
Default branch
main
View the original SKILL.md

OpenFold3 Structure Prediction

Prerequisites

RequirementMinimumRecommended
Python3.10+3.11
CUDA12.1+12.4+
GPU VRAM24GB80GB (H100)
RAM32GB64GB
Disk (weights)3GB-

How to run

Installation

pip install 'openfold3[cuequivariance]==0.4.1'

The default attention kernel is DeepSpeed DS4Sci_EvoformerAttention. If DeepSpeed is unavailable, switch to the cuEquivariance triangle kernels (no build-from-source) by overriding the eval memory settings in model_config.py (use_deepspeed_evo_attention: False, use_cueq_triangle_kernels: True). Some pre-built environments already ship this override; check before re-patching.

Weights

Apache-2.0, ~2.3 GB from HF OpenFold/OpenFold3. The repo is gated (auto-approval) — accept the access form on the HF model page and authenticate (huggingface-cli login or HF_TOKEN) before downloading:

export OPENFOLD_CACHE=~/.openfold3
huggingface-cli download OpenFold/OpenFold3 checkpoints/of3-p2-155k.pt \
  --local-dir "$OPENFOLD_CACHE"

run_openfold will also auto-download to $OPENFOLD_CACHE on first run if egress is open and HF credentials are available (either HF_TOKEN or a prior huggingface-cli login) with repo access granted. The interactive setup_openfold helper exists but prompts on stdin; prefer the explicit download above for non-interactive runs.

Running

export OPENFOLD_CACHE=/path/to/cache
run_openfold predict \
  --query_json=queries.json \
  --output-dir out/ \
  --use-msa-server false \
  --use-templates false

run_openfold discovers the checkpoint under $OPENFOLD_CACHE automatically. Only pass --inference-ckpt-path <file.pt> if you have a non-standard layout or multiple checkpoints and need to pin one explicitly.

For MSA + templates (slower, higher accuracy), drop the two false flags. The MSA server is api.colabfold.com; template chain-ID remap hits data.rcsb.org (GraphQL) — both must be reachable.

Query JSON format

OpenFold3 does not read FASTA. Queries are a JSON object validated by InferenceQuerySet (pydantic, extra: forbid — unknown keys reject):

{
  "queries": {
    "my_complex": {
      "chains": [
        {"molecule_type": "protein", "chain_ids": ["A"], "sequence": "MQIFVK…"},
        {"molecule_type": "protein", "chain_ids": ["B", "C"], "sequence": "MVLSPA…"},
        {"molecule_type": "ligand",  "chain_ids": ["L"], "smiles": "CC(=O)Oc1ccccc1C(=O)O"}
      ],
      "use_msas": true
    }
  },
  "seeds": [42]
}
molecule_typerequired field
protein / dna / rnasequence
ligandsmiles or ccd_codes: ["HEM"]

chain_ids is a list — repeat the same sequence across multiple chain IDs for homo-oligomers. Per-chain paired_msa_file_paths / main_msa_file_paths let you supply your own a3m instead of the server.

Key parameters

FlagDefaultDescription
--num-diffusion-samples5Structures per (query, seed)
--num-model-seeds1Number of model seeds per query (multiplies output count alongside JSON seeds and diffusion samples)
--use-msa-servertrueColabFold MMseqs2 server for MSA
--use-templatestrueColabFold template search + RCSB remap
--inference-ckpt-pathauto-discovered under $OPENFOLD_CACHEOverride only — for non-standard layouts or to pin a specific checkpoint file

Wisp execution

Use python only for bounded interactive checks. For a long or GPU-backed workload, require a selected and probed ssh:<alias> context and load remote-compute-ssh. Put the documented invocation in a self-contained project script, activate the remote environment explicitly, stage only small files with input_paths, and make the command write to a known absolute remote result path. Submit it with run_in_context and register that exact ssh:// path in output_specs. Call monitor_run once when waiting is needed, get_run once for a snapshot, or cancel_run to stop. Do not send a scheduler submission through the SSH-direct runner.

Output format

out/
├── summary.txt
├── model_config.json / experiment_config.json
├── inference_query_set.json
└── <query_name>/seed_<N>/
    ├── <query>_seed_<N>_sample_<k>_model.cif
    ├── <query>_seed_<N>_sample_<k>_confidences.json           # full PAE/pLDDT
    ├── <query>_seed_<N>_sample_<k>_confidences_aggregated.json
    └── timing.json

*_confidences_aggregated.json is the small one to read first:

{
  "avg_plddt": 78.96, "ptm": 0.667, "iptm": 0.0, "gpde": 0.73,
  "has_clash": 0.0, "sample_ranking_score": 0.133,
  "chain_ptm": {"A": 0.667}, "chain_pair_iptm": {}
}

What good output looks like

  • summary.txt shows Successful Queries: N matching your input count
  • avg_plddt > 70 (single-seq) / > 80 (with MSA)
  • ptm > 0.6; for complexes, iptm > 0.5
  • has_clash: 0.0
  • .cif ~50-150 KB per sample for a small protein

Verify

grep -E 'Successful|Failed' out/summary.txt
find out -name '*_model.cif' | wc -l   # = queries x json_seeds x num-model-seeds x num-diffusion-samples

Troubleshooting

ErrorCauseFix
_deepspeed_evo_attn requires that DeepSpeed be installeddefault eval kernel is DS4Sci on CUDAinstall deepspeed (needs nvcc + CUTLASS), or in model_config.py eval block set use_deepspeed_evo_attention: False + use_cueq_triangle_kernels: True (cuEq path; no build)
CUTLASS_PATH ... not set ... cutlass_library is not installedcuEq path still needs the python cutlass_library shimpip install nvidia-cutlass
libXrender.so.1: cannot open shared object filerdkit (via pdbeccdutils) needs X11 render libsapt-get install libxrender1 libxext6 libsm6
ModuleNotFoundError: boto3 (or awscrt)openfold3.core.data.io.s3 is eager-imported even when weights are localpip install boto3 awscrt
ValidationError: queries / Field required or Input should be an objectwrong JSON shapetop-level is {"queries": {"<name>": {...}}} (a dict, not a list)
ValidationError ... settings / Extra inputs are not permittedtried to override model config via --runner-yaml--runner-yaml is InferenceExperimentConfig only; kernel/memory settings live in model_config.py
Failed to fetch chain ID mappings from RCSB for N entriesdata.rcsb.org unreachable (allowlist/offline)run with --use-templates false, or open egress to data.rcsb.org
CUDA out of memorylarge complex / many samplesreduce --num-diffusion-samples; the low_mem preset (model_setting_presets.yml) offloads more aggressively