La-Ribo

RNA Co-Design via Geometry–Latent Flow Matching

Generate RNA sequences and three-dimensional heavy-atom structures together.

Paper Code Weights License

Quickstart · Models · Files · Dataset · Citation


La-Ribo jointly samples a phosphate–sugar–base scaffold and residue-wise latent vectors. A shared autoencoder decodes this representation into nucleotide identities and heavy-atom coordinates.

This repository provides two pretrained flow models, their shared autoencoder and normalization statistics, and 10,631 trRosettaRNA2-predicted RNA structures released with the paper.

Models

La-Ribo-Base La-Ribo-Tri
Pair-feature updates No triangle updates Triangle updates every 2 blocks
Flow layers / latent dimension 12 / 16 12 / 16
Flow checkpoint la_ribo_base.pt la_ribo_tri.pt
Config in the code repository config/la-ribo-base.yaml config/la-ribo-tri.yaml

Both variants use EMA flow weights, the same VAE, and the same statistics files. The default sampling settings are 200 Euler steps, scaffold power 1.25, and latent power 2.0.

Quickstart

1. Install the inference code

Install uv, then:

git clone https://github.com/GENTEL-lab/La-Ribo.git
cd La-Ribo
uv python install 3.14
uv sync --locked

Requirements: Python 3.14+; CPU, including macOS, or NVIDIA CUDA on Linux. CUDA inference requires a CUDA 13-compatible driver.

2. Download the model bundle

Run this from the cloned code directory. The dataset is optional and is not included in this command.

uvx --from huggingface_hub hf download GENTEL-Lab/La-Ribo \
  config.json \
  weights/la_ribo_base.pt \
  weights/la_ribo_tri.pt \
  weights/la_ribo_vae.pt \
  weights/latent_stats.pt \
  weights/coarse_stats.pt \
  --local-dir .

3. Generate with La-Ribo-Tri

uv run --no-sync python sample_latent_flow.py \
  --config config/la-ribo-tri.yaml \
  --checkpoint weights/la_ribo_tri.pt \
  --autoencoder_checkpoint weights/la_ribo_vae.pt \
  --latent_stats weights/latent_stats.pt \
  --coarse_stats weights/coarse_stats.pt \
  --num_res 80 \
  --seed 2027 \
  --device auto \
  --output samples/tri_80_seed2027.pdb
Generate with La-Ribo-Base
uv run --no-sync python sample_latent_flow.py \
  --config config/la-ribo-base.yaml \
  --checkpoint weights/la_ribo_base.pt \
  --autoencoder_checkpoint weights/la_ribo_vae.pt \
  --latent_stats weights/latent_stats.pt \
  --coarse_stats weights/coarse_stats.pt \
  --num_res 80 \
  --seed 2027 \
  --device auto \
  --output samples/base_80_seed2027.pdb
To change… Set…
RNA length --num_res; the paper benchmark covers 40–150 nt
Random sample --seed, together with a new output filename
Compute device --device cpu or --device cuda; auto selects CUDA when available, otherwise CPU
NVIDIA GPU Prefix the command with CUDA_VISIBLE_DEVICES=0, replacing 0 with the desired index

Outputs: a PDB containing the generated sequence and heavy-atom coordinates, plus a neighboring .summary.json with sampling parameters and provenance. Output directories are created automatically; existing samples are not overwritten.

These commands perform sampling. Refolding and designability evaluation are separate procedures. Matching CPU/CUDA seeds alone does not guarantee identical structures because the backends use different random-number implementations.

Release files

File Contents
config.json Variant definitions, architecture dimensions, file locations, and sampling defaults
weights/la_ribo_base.pt Base EMA state dictionary
weights/la_ribo_tri.pt Tri EMA state dictionary
weights/la_ribo_vae.pt Shared autoencoder state dictionary
weights/latent_stats.pt Latent whitening statistics
weights/coarse_stats.pt Scaffold scaling statistics

The model files are pure tensor state dictionaries. EMA is already applied to the flow weights, so no --use_ema flag is required. Both statistics files are needed to reconstruct the correct latent and coordinate scales; they contain only inference-required values and format fields.

Dataset

The data/ directory contains 10,631 RNA sequences, corresponding trRosettaRNA2-predicted structures, and per-residue pLDDT scores. It is optional for pretrained-model inference.

data/
├── metadata.csv
└── trran_data_publish.tar.gz
    └── <name>/
        ├── model_1_relaxed200.pdb
        └── plddt.csv
uvx --from huggingface_hub hf download GENTEL-Lab/La-Ribo \
  --include "data/*" --local-dir .
tar -xzf data/trran_data_publish.tar.gz -C data
Dataset metadata and archive format

metadata.csv contains one row per RNA:

Column Description
name Entry ID and archive directory name, e.g. 000011_bpRNA_RFAM_7168_L89
original_task_id Task ID from the prediction pipeline
source_db bpRNA-1m90 (6,458) or RNAStrAlign (4,173)
source_id ID in the source database
length Sequence length: 16–252 nt, median 77
mean_plddt Mean pLDDT, on a 0–1 scale: 0.80–0.99
min_residue_plddt Minimum per-residue pLDDT
max_residue_plddt Maximum per-residue pLDDT
cluster_rep_id90_cov95 Sequence-cluster representative: 10,139 clusters
sequence RNA sequence

The approximately 539 MB archive extracts to trran_data_publish/. Each entry directory contains model_1_relaxed200.pdb and plddt.csv; the CSV columns are Residue_Index and plddt.

Download statistics

The model-bundle command includes config.json, which Hugging Face uses as this repository's default download-counting signal. Downloading only .pt files or files under data/ does not trigger that default model counter.

HF counts matching GET/HEAD requests server-side. The displayed metric is not a count of unique users or guaranteed completed weight transfers. How model downloads are counted →

Citation

@article{ma2026laribo,
  title   = {La-Ribo: RNA Co-Design via Geometry-Latent Flow Matching},
  author  = {Ma, Runze and Hua, Will and Zheng, Shuangjia},
  year    = {2026},
  eprint  = {2610.12236},
  archivePrefix = {arXiv},
  primaryClass  = {q-bio.BM},
  url     = {https://arxiv.org/abs/2610.12236}
}

Released under the MIT License. See the code repository for implementation details.

Downloads last month
13
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for GENTEL-Lab/La-Ribo