La-Ribo
RNA Co-Design via Geometry–Latent Flow Matching
Generate RNA sequences and three-dimensional heavy-atom structures together.
Quickstart · Models · Files · Dataset · Citation
La-Ribo jointly samples a phosphate–sugar–base scaffold and residue-wise latent vectors. A shared autoencoder decodes this representation into nucleotide identities and heavy-atom coordinates.
This repository provides two pretrained flow models, their shared autoencoder and normalization statistics, and 10,631 trRosettaRNA2-predicted RNA structures released with the paper.
Models
| La-Ribo-Base | La-Ribo-Tri | |
|---|---|---|
| Pair-feature updates | No triangle updates | Triangle updates every 2 blocks |
| Flow layers / latent dimension | 12 / 16 | 12 / 16 |
| Flow checkpoint | la_ribo_base.pt |
la_ribo_tri.pt |
| Config in the code repository | config/la-ribo-base.yaml |
config/la-ribo-tri.yaml |
Both variants use EMA flow weights, the same VAE, and the same statistics files. The default sampling settings are 200 Euler steps, scaffold power 1.25, and latent power 2.0.
Quickstart
1. Install the inference code
Install uv, then:
git clone https://github.com/GENTEL-lab/La-Ribo.git
cd La-Ribo
uv python install 3.14
uv sync --locked
Requirements: Python 3.14+; CPU, including macOS, or NVIDIA CUDA on Linux. CUDA inference requires a CUDA 13-compatible driver.
2. Download the model bundle
Run this from the cloned code directory. The dataset is optional and is not included in this command.
uvx --from huggingface_hub hf download GENTEL-Lab/La-Ribo \
config.json \
weights/la_ribo_base.pt \
weights/la_ribo_tri.pt \
weights/la_ribo_vae.pt \
weights/latent_stats.pt \
weights/coarse_stats.pt \
--local-dir .
3. Generate with La-Ribo-Tri
uv run --no-sync python sample_latent_flow.py \
--config config/la-ribo-tri.yaml \
--checkpoint weights/la_ribo_tri.pt \
--autoencoder_checkpoint weights/la_ribo_vae.pt \
--latent_stats weights/latent_stats.pt \
--coarse_stats weights/coarse_stats.pt \
--num_res 80 \
--seed 2027 \
--device auto \
--output samples/tri_80_seed2027.pdb
Generate with La-Ribo-Base
uv run --no-sync python sample_latent_flow.py \
--config config/la-ribo-base.yaml \
--checkpoint weights/la_ribo_base.pt \
--autoencoder_checkpoint weights/la_ribo_vae.pt \
--latent_stats weights/latent_stats.pt \
--coarse_stats weights/coarse_stats.pt \
--num_res 80 \
--seed 2027 \
--device auto \
--output samples/base_80_seed2027.pdb
| To change… | Set… |
|---|---|
| RNA length | --num_res; the paper benchmark covers 40–150 nt |
| Random sample | --seed, together with a new output filename |
| Compute device | --device cpu or --device cuda; auto selects CUDA when available, otherwise CPU |
| NVIDIA GPU | Prefix the command with CUDA_VISIBLE_DEVICES=0, replacing 0 with the desired index |
Outputs: a PDB containing the generated sequence and heavy-atom coordinates, plus a neighboring .summary.json with sampling parameters and provenance. Output directories are created automatically; existing samples are not overwritten.
These commands perform sampling. Refolding and designability evaluation are separate procedures. Matching CPU/CUDA seeds alone does not guarantee identical structures because the backends use different random-number implementations.
Release files
| File | Contents |
|---|---|
config.json |
Variant definitions, architecture dimensions, file locations, and sampling defaults |
weights/la_ribo_base.pt |
Base EMA state dictionary |
weights/la_ribo_tri.pt |
Tri EMA state dictionary |
weights/la_ribo_vae.pt |
Shared autoencoder state dictionary |
weights/latent_stats.pt |
Latent whitening statistics |
weights/coarse_stats.pt |
Scaffold scaling statistics |
The model files are pure tensor state dictionaries. EMA is already applied to the flow weights, so no --use_ema flag is required. Both statistics files are needed to reconstruct the correct latent and coordinate scales; they contain only inference-required values and format fields.
Dataset
The data/ directory contains 10,631 RNA sequences, corresponding trRosettaRNA2-predicted structures, and per-residue pLDDT scores. It is optional for pretrained-model inference.
data/
├── metadata.csv
└── trran_data_publish.tar.gz
└── <name>/
├── model_1_relaxed200.pdb
└── plddt.csv
uvx --from huggingface_hub hf download GENTEL-Lab/La-Ribo \
--include "data/*" --local-dir .
tar -xzf data/trran_data_publish.tar.gz -C data
Dataset metadata and archive format
metadata.csv contains one row per RNA:
| Column | Description |
|---|---|
name |
Entry ID and archive directory name, e.g. 000011_bpRNA_RFAM_7168_L89 |
original_task_id |
Task ID from the prediction pipeline |
source_db |
bpRNA-1m90 (6,458) or RNAStrAlign (4,173) |
source_id |
ID in the source database |
length |
Sequence length: 16–252 nt, median 77 |
mean_plddt |
Mean pLDDT, on a 0–1 scale: 0.80–0.99 |
min_residue_plddt |
Minimum per-residue pLDDT |
max_residue_plddt |
Maximum per-residue pLDDT |
cluster_rep_id90_cov95 |
Sequence-cluster representative: 10,139 clusters |
sequence |
RNA sequence |
The approximately 539 MB archive extracts to trran_data_publish/. Each entry directory contains model_1_relaxed200.pdb and plddt.csv; the CSV columns are Residue_Index and plddt.
Download statistics
The model-bundle command includes config.json, which Hugging Face uses as this repository's default download-counting signal. Downloading only .pt files or files under data/ does not trigger that default model counter.
HF counts matching GET/HEAD requests server-side. The displayed metric is not a count of unique users or guaranteed completed weight transfers. How model downloads are counted →
Citation
@article{ma2026laribo,
title = {La-Ribo: RNA Co-Design via Geometry-Latent Flow Matching},
author = {Ma, Runze and Hua, Will and Zheng, Shuangjia},
year = {2026},
eprint = {2610.12236},
archivePrefix = {arXiv},
primaryClass = {q-bio.BM},
url = {https://arxiv.org/abs/2610.12236}
}
Released under the MIT License. See the code repository for implementation details.
- Downloads last month
- 13