FullTilt

NeurIPS 2026Main Track

At FullTilt Real-Time Open-Set 3D Macromolecule Detection Directly from Tilted 2D Projections

Duke University

The first end-to-end cryo‑ET 3D detector that works directly on the tilt-series.

No tomogram inference, no sliding windows: all tilt images go in together, and 3D particle coordinates come out in under a second.

Raw aligned tilt image of EMPIAR-10304 tilt1
Input tilt-series
Ground-truth ribosome locations projected onto the tilt image
Ground truth
FullTilt 3D detections projected onto the tilt image
FullTilt · 1 prompt
0°

EMPIAR-10304 · tilt1, purified E. coli 70S ribosomes. Scrub or play the tilt-series. FullTilt reads all 41 tilt images at once and predicts 3D particle centers; we project them back onto each tilt image. The blue box is the single visual prompt. Use ← → to step, space to play or pause.

0.23s
for a full tilt-series on EMPIAR-10304
>7,490×
faster than TomoTwin, with higher F1
8.6×
less peak VRAM (2.4 GB vs 20.8 GB)
0
retraining for new targets — open-set, zero-shot

01Abstract

Open-set 3D macromolecule detection in cryogenic electron tomography eliminates the need for target-specific model retraining. However, strict VRAM constraints prohibit processing an entire 3D tomogram, forcing current methods to rely on slow sliding-window inference over extracted subvolumes.

To overcome this, we propose FullTilt, an end-to-end framework that redefines 3D detection by operating directly on aligned 2D tilt-series. Because a tilt-series contains significantly fewer images than slices in a reconstructed tomogram, FullTilt eliminates redundant volumetric computation, accelerating inference by orders of magnitude. To process the entire tilt-series simultaneously, we introduce a tilt-series encoder to efficiently fuse cross-view information. We further propose a multiclass visual prompt encoder for flexible prompting, a tilt-aware query initializer to effectively anchor 3D queries, and an auxiliary geometric primitives module to enhance the model's understanding of multi-view geometry while improving robustness to adverse imaging artifacts. Extensive evaluations on three real-world datasets demonstrate that FullTilt achieves state-of-the-art zero-shot performance while drastically reducing runtime and VRAM requirements, paving the way for rapid, large-scale visual proteomics analysis.

02Skip the tomogram

A tomogram is reconstructed from the aligned tilt-series, so the tilt-series already holds at least as much information — in far fewer images (e.g. 41 tilts versus 256 depth slices). Existing open-set pickers still work on the reconstructed volume, sliding a model over thousands of subvolumes or slices. FullTilt instead maps the 2D tilt-series to 3D detections in one forward pass.

Four detection pipelines: 3D models on tomogram subvolumes, 2D models on tomogram slices, 2D detection with back-projection, and FullTilt's end-to-end 2D-to-3D model.
Comparison of detection mechanisms. (a) TomoTwin and ProPicker slide 3D models over subvolumes; (b) CryoSAM runs SAM over orthogonal slices; (c) Zeng et al. detect in each tilt image and back-project; (d) FullTilt processes the whole tilt-series end-to-end.

The catch: cryo-ET projections have an extremely low signal-to-noise ratio. A particle is often only recognizable across consecutive tilts, particles overlap because each image projects through the full sample depth, and high tilts suffer occlusion and uneven illumination. Per-image detectors and off-the-shelf multi-view 3D detectors (DETR3D, PETR) break down here.

Three tilt images at −45°, 0° and 45° with thyroglobulin particles boxed, next to one tomogram slice.
2D tilt-series vs. 3D tomogram (CZII, TS_6_6). In single tilt images, particles are noisy, overlapping and sometimes occluded (red arrow); in a tomogram slice they are separated in depth.

03Method

FullTilt is a DETR-style encoder–decoder. A 2D backbone encodes every tilt image; four new components turn those 2D features into 3D detections of whatever the user prompts for.

FullTilt architecture: 2D backbone, tilt-series encoder, multiclass visual prompt encoder, tilt-aware query initializer and 3D decoder.
FullTilt architecture. (a) Overview. (b) Tilt-series encoder. (c) Multiclass visual prompt encoder. (d) 3D decoder.

ℰTiltTilt-series encoder

Alternates local attention within each image with global row attention across tilts. Because the series is aligned to a vertical tilt axis, a particle stays on the same row in every image — so cross-view attention only needs to look along rows.

ℰPromptMulticlass visual prompt encoder

Turns any number of 2D box prompts, for any number of classes, into one prototype per class using masked attention. It handles images with no prompts at all, and plugs into any DETR-like detector.

ℐTiltTilt-aware query initializer

At a 0° tilt a particle's projection gives its x and y directly. FullTilt picks 3D anchors from the prototype-matched peaks of the near-zero tilt, starting depth at the mid-plane.

𝒢Auxiliary geometric primitives

Generates synthetic projections of simple shapes with exact 3D labels on the fly, including dense clusters, occlusion and uneven illumination, to teach multi-view geometry cheaply.

04Results

Zero-shot, intra-instance detection on three real-world cryo-ET datasets. Prompts and targets come from the same tomogram; results are the mean ± std over 10 trials. Runtime and peak VRAM are measured on one NVIDIA RTX 6000 Ada.

Prompts

F1 vs. runtime

Up and to the left is better. Runtime on a log scale.

Method mAP@0.5r ↑ mAP@1r ↑ F1 ↑ Runtime (s) ↓ VRAM (GB) ↓

Bold marks the best cryo-ET method per column. †Equipped with our multiclass visual prompt encoder ℰPrompt. n.a.: the method gives no confidence scores, so mAP cannot be computed. Cross-instance, multiclass and ablation results are in the paper.

05In the tomogram

FullTilt never sees the tomogram, but its 3D predictions land on the particles when drawn on reconstructed Z-slices. The same EMPIAR-10304 tomogram, single prompt, boxes shown within ±r of each slice.

FullTilt detections on tomogram slices z=112, 118 and 124
FullTilt · 0.23 s, from the tilt-series alone.

06Limitations

FullTilt struggles with small macromolecules: AP is near zero for every evaluated protein below 100 kDa, in line with the small-object difficulties of DETR-like detectors. Zero-shot performance is still below fully supervised methods, likely due to the gap between simulated training data and real tomograms, which motivates real-world fine-tuning and larger-scale training with FullTilt as a foundation.

07BibTeX

@inproceedings{ho2026fulltilt,
  title     = {At FullTilt: Real-Time Open-Set 3D Macromolecule Detection
               Directly from Tilted 2D Projections},
  author    = {Ho, Ming-Yang and Bartesaghi, Alberto},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026}
}