For the complete documentation index, see llms.txt. This page is also available as Markdown.

Advanced automated data processing: a case study using CAK

Introduction

Automated cryo-EM data processing is an important area of development because it can reduce time-to-structure, increase the number of parallel datasets that can be processed, lower the barrier to entry for new practitioners and ensure reproducible results across experiments.

We initially demonstrated an automation strategy with a test set of 21 challenging GPCR datasets, processed fully automatically, using a general automation Workflow and strategy that can be easily adapted to new target classes (”General Automation Workflow v1”). The Workflow and tools that were developed in CryoSPARC to support automated data processing were able to produce results equal or better than manual processing with zero manual intervention for repeat target scenarios.

To further demonstrate generalizable automation, and to serve as a guide for users applying CryoSPARC automation to their own targets, this case study covers the application of automation to a series of nine drug-bound CDK-activating Kinase (CAK) datasets. CAK is a very small, soluble, and druggable protein complex that is an oncogenic and antiviral target of interest due to its central role in physiology. Our efforts to automate data processing build upon the original CAK dataset depositors' (Cushing et al., 2024) goal of establishing high-throughput automated data collection, and together these improvements can help accelerate drug discovery workflows, especially for small targets like CAK.

We provide an updated and improved version of a general automation Workflow, “General Automation Workflow v2”, and show that it produces equal or better results on CAK datasets compared to manual processing. The updated Workflow is more robust to preferred orientation than the previous version and can be downloaded and used to automate processing of your own target.

The case study is organized as follows:

Background on CDK-activating Kinase (CAK) and Description of Test Datasets

CAK (CDK-Activating Kinase) is a critical regulator of cell division and growth. This complex consists of three subunits—MAT1, Cyclin H, and CDK7—wherein ligands bind specifically to a cleft in CDK7 (Figure 1, green ligand). Despite being a significant target for antiviral and cancer therapies, CAK remains a notoriously challenging cryo-EM subject due to its small size (approximately 80 kDa).

Figure 1. Cartoon of the CAK complex highlighting many important features. Low resolution surface representations of the CAK complex including CDK7 (blue), Cyclin H (grey), MAT1 (orange), and the ligand (green). All ligands bind in the same cleft as the ATPγS molecule depicted.

This case study uses nine EMPIAR datasets, encompassing CAK in its apo state, bound to a nucleotide analog, and in complex with seven distinct inhibitors (Table 1). The datasets featured in this case study were originally processed by the depositors using a manual approach involving CryoSPARC (versions 3.3.1-4.1.1) and RELION 4.0. To give a brief overview of the depositors data processing approach, micrograph preprocessing were completed during data collection within CryoSPARC Live. Particles were then picked using the Blob Picker and subjected to 2D Classification; ultimately, particles from the selected classes were reconstructed in CryoSPARC Live using Homogeneous Refinement to yield a starting volume. Once data collection was finalized, the depositors utilized templates from the CryoSPARC Live Streaming 2D Classification, as well as templates generated from the volume output of the CryoSPARC Live Homogeneous Refinement, to perform template picking across the entire dataset in separate jobs. All blob and template-picked particles underwent duplicate removal before being converted to the STAR file format and exported to RELION. Subsequently, a series of 3D auto-refinements, 3D classifications, global CTF refinements, and Bayesian particle polishing jobs were executed to achieve the high-resolution reconstruction utilized for model building. While this represents the general workflow employed by the depositors, they noted that certain targets and samples necessitated additional, specialized steps.

EMPIAR
EMDB
PDB
Inhibitor
Grid name
# of movies
Final # of particles yielding deposited map

11800

17523

---

Apo

VC15-1

5,805

486K

11793

17129

8ORM

THZ1

29-1

9,118

555K

11799

17511

8P6Y

ATPgS

VC16-3

5,708

126k

11807

17754

8PLZ

CT7030

BG51-1

10,017

177K

11823a

17520

8P77

ICE-C0943

VC1-1

5,729

294k

11823b

17515

8P72

ICE-C0768

VC2-4

5,583

335K

11823c

17521

8P78

Dinaciclib

VC4-3

5,210

295K

11823d

17509

8P6W

BS181

VC8-1

5,203

565K

11823ef

17508

8P6V

ICE-C0942(a)

VC13-1/VC14-1

10,598

270K

Table 1. CAK Automation Dataset Information - Information for all nine datasets used in this study including the EMPIAR, EMDB, and PDB depositions. All datasets were collected on a Krios G4 operating at 300 keV and equipped with a Falcon 4i and SelectrisX energy filter. All EER formatted movies were collected using the same data collection parameters: pixel size of 0.57Å, a dose of 70e-/Å2, 70 fractions, and on a microscope operating Cs value of 2.7, and at 300 keV. Dataset size ~2.2 TB/5,000 movies.

The CAK datasets described above are challenging to process for a few reasons. The CAK complex is small (~85 kDa) and therefore the signal-to-noise (SNR) ratio, which affects both particle picking and alignment, is limited and drops quickly in thicker ice. While the depositing authors were able to make grids with very thin ice and therefore collect datasets with improved SNR, the grids presented significant preferred orientation which hindered the final quality of the reconstructions. Looking below to Figure 2, although we can see that while high resolution features are present, including holes in some aromatic rings, there is stretching and fragmentation of the density across the map in the direction of the poorly resolved orientation. Global orientation metrics for the deposited map, such as the low cFAR score and the missing wedge of information in the 3DFSC central slice, highlight the challenges presented by these datasets. Many cryo-EM datasets have similar issues, and preferred orientation is a common reason for failure of cryo-EM reconstruction.

Figure 2. Anisotropy present in the deposited CAK maps presents a challenging case to deal with preferred orientation in an automated manner. (A) Face and side views of the deposited map for EMPIAR-11823d highlighting the high resolution features, but also the presence of the anisotropic stretching. (B) cFAR plot for the deposited map. (C) Plot of all 3072 cFSCs showing the wide range of resolutions across the viewing sphere. (D) Central slices of 3DFSC volume showing the missing wedge of information due to preferred orientation.

An Updated Automation Workflow: Workflow v2

The general automation strategy that we use for fully automated data processing (described in detail in this preprint) involves standard preprocessing, high-quality generalizable particle picking, extensive and robust particle curation, and multiple refinements to achieve the highest quality results. In our initial demonstration of this strategy for repeat GPCR structure determination, we developed a CryoSPARC Workflow (Automation Workflow v1) that could be easily generalized to different target classes by simply providing a new target-class-specific low-resolution (15Å) reference volume.

For CAK, we used Workflow v1 as a starting point for processing a single apo-CAK dataset (EMPIAR-11800), and found that it produced results similar to the deposited maps, but ran into the challenges outlined above, in particular, strong preferred orientation present in the data. We therefore modified and updated the automation workflow, following the same overall strategy but simplifying particle picking and adding additional steps to account for preferred orientation. The experimental development leading to this updated workflow, Workflow v2, is detailed in a section below.

Workflow v2 (Figure 3) provides simplified, robust, and fully automated end-to-end processing. Thanks to the general automation strategy on which it is based, Workflow v2 can be run automatically on new datasets with no changes (other than dataset-level inputs such as microscope parameters). It can also be generalized to new target classes by changing only the target-class/workflow level inputs (a low resolution reference volume, particle diameters, thresholds, and junk volumes).

Figure 3. Fully-automated workflow (Workflow v2) developed for processing all CAK datasets that relies on modified picking and curation strategies to reduce effects of preferred orientation. This figure shows a majority of jobs in the workflow along with all required inputs for those jobs, and it can be easily adapted to new targets with a few key adjustments to the highlighted parameters.

As shown in Figure 3, Workflow v2 contains all end-to-end data processing steps. Preprocessing steps including Patch Motion Correction and Patch CTF Estimation followed by the Micrograph Denoiser and Micrograph Junk Detector. Using the Template Picker, particles are picked using the templates created from the input low resolution reference volume and a user supplied particle diameter and separation distance. Particle picks are then cleaned using the auto option in Inspect Particle Picks to remove particle picks on the background and picks on junk regions are removed using the Micrograph Junk Detector.

After particle picks have been cleaned, particles are extracted and submitted to decoy classification — Heterogeneous Refinement using as inputs the cleaned particle stack, the low resolution reference volume, and four junk volumes to attract poor particles. Note that the junk volumes we use for CAK are identical to those used for GPCR datasets in Workflow v1. The good volume and particle stack are then subjected to a Non-uniform Refinement to improve particle poses and compute per-particle scale values. Rebalance Orientations is then used to filter out particles in over-represented viewing directions. A second round of particle curation is then performed on the rebalanced particle stack where a three-class Ab-Initio Reconstruction is used followed by Heterogeneous Refinement. Using Reference Based Auto Select 3D, the highest resolution volume(s) and particle stack(s) are selected for additional pruning of over-represented views using Rebalance Orientations followed by removal of low scale particles with Subset Particles by Statistic. Together, these additional stages greatly reduce map anisotropy by pruning over-represented views and removing low-quality particles.

memo-pad

Note: Workflow v2 makes use of Rebalance Orientations to prune heavily overrepresented views. If you are using Workflow v2 on a target that exhibits little or no preferred orientation, the Rebalance percentile can be increased to a value of 80-90 so as to preserve more of the picked particles.

For final refinements, particles are subjected to Non-uniform Refinement, Reference Based Motion Correction, and a final Non-uniform Refinement. Here, on-the-fly global CTF refinement in the Non-uniform Refinement job is used to correct for residual beam tilt left over from AFIS data collection. Figure 4, below, details the differences between the v1 and v2 versions of the workflow.

Figure 4. Changes made between Workflow v1 and Workflow v2. Workflow v1 was originally developed for general automation and applied to a test set of 21 GPCRs and Workflow v2 is an improved iteration applied to the CAK datasets in this case study. The overall structure of the Workflows is similar; the picking strategy was changed in Workflow v2 and we introduced Rebalance Orientations and Subset Particles by Statistic to increase robustness to severe preferred orientation.

In the next section, we show results of automated processing of CAK datasets. To confirm the generalizability of Workflow v2, we also used it to reprocess GPCR datasets and found that it produces similar or improved results compared to the original Workflow (v1).

Fully Automated Processing Results for CAK

Using Workflow v2, we processed all eight ligand-bound CAK datasets along with the apo dataset in a fully hands-off manner. Below we summarize map metrics and improvements to map quality and the resulting models that stem from a higher quality reconstruction.

Data and Input Preparation

We processed the apo-CAK dataset manually by following the stages of Workflow v1 (see below) and the resulting refined volume, low-pass filtered to 15Å, became the reference volume used for the other eight CAK datasets. The reference volume is used in two places in the automated workflow (Figure 3), to create templates for template picking, and as the good volume during decoy classification.

Automated processing requires setting dataset-level, target-class-level and workflow-level inputs. Dataset names used below are based on their EMPIAR ID, and for EMPIAR-11823, the names are assigned alphabetically (a-f) based on the order of the datasets in the deposition. Full dataset information can be found in Table 1. Dataset-specific microscope parameters for each dataset were obtained from the respective EMPIAR depositions and these values were the same for all datasets: 0.57 Å pixel size, 70 electrons/Å2, and EER format fractionated into 70 frames. Next, based on processing the apo-CAK dataset, multiple class and workflow level parameters were selected. For particle picking, a particle diameter of 90Å and a separation distance of 0.6 diameters (54Å) were selected. For initially cleaning particle picks, it was necessary to adjust the power score of the Inspect Particle Picks auto-cluster mode to 20 since the CAK complex is soluble (lacks a micelle) and is much smaller than a GPCR. The final refinement map of apo-CAK was lowpass filtered to 15Å and used in conjunction with the GPCR junk volumes for decoy classification. Unlike processing the GPCR datasets where the class-level parameters required changes, these target datasets are all extremely similar and thus none of these parameters were varied between them.

FSC Resolution and Map Anisotropy

To be useful in practice, automated data processing should produce results equal or better than manual processing in terms of map quality and downstream interpretability. The first measures of success of automated data processing of the CAK complex datasets were the resolution and cFAR scores of the final volumes; summarized below in Figure 5. Comparing deposited manual processing results to our automated results, the resolution of the apo dataset was improved by 0.1Å, was unchanged for EMPIAR 11799, and was within 0.1Å to 0.3Å for all other datasets. Since global resolutions do not always provide a complete picture for datasets that have preferred orientation, we also examined map anisotropy. cFAR scores, show a major improvement in the automated results: values are near-zero for most deposited maps, indicating poor map quality due to preferred orientation, versus an average increase of 0.33 for automated maps, indicating interpretable and higher quality map density.

Figure 5. Global map metrics confirm that automated maps have similar resolution but significantly improved orientation distributions compared to deposited maps. Left: Gold-standard Fourier Shell Correlation (FSC) and right: Conical FSC Area Ratio (cFAR). Values from deposited published results shown as gray circles. Values from automated processing shown as filled, colored circles: improved (green), equivalent (blue), or reduced (orange).

Improved Map Quality and Interpretability

To further evaluate the results of automated processing, we examined visual map quality and how well the ligand was resolved. In Figure 6, we have summarized the results of the deposited, manually processed data alongside our results from automated data processing. Map pairs are shown at the same number of standard deviations from the mean (9) for equivalent comparison. All maps, both deposited and automated, contain sidechain density, and for some maps, holes in large aromatic residues are present. In all cases, all three protein chains for CDK7, Cyclin H, and MAT1 are well resolved, and the ligand is present in the density.

For most of the deposited maps, the face views do not appear to exhibit anisotropy, but in the side views, there is strong anisotropy that manifests as stretching of the density along with increased discontinuity in regions, especially density for the protein backbone. Video 1 shows this phenomenon for EMPIAR-11823d where upon rotation of the volume, the comparison is quite striking. Comparing the deposited maps to the automated maps, the map anisotropy is greatly reduced in the automated results and these observations of anisotropy are confirmed by the improvement in the cFAR scores in Figure 5.

Figure 6. Automated maps from the final Non-uniform Refinement exhibit well resolved density and less anisotropic effects compared to the deposited maps. Deposited maps and our automated processing maps are shown at similar thresholds for all datasets. Map coloring follows the same color scheme as Figure 1. For each global map, FSC resolutions (left, boxed) and cFAR score (left, unboxed) are shown. For automated maps, color scheme denotes improved (green), equivalent (blue), or worse (orange) relative to the manually processed deposited maps.
Video 1. Map density from automated processing is of higher quality as compared with the deposited density. This video shows the reduction of the map anisotropy for the automated processing of EMPIAR-11823d (left) compared to the deposited map (right).
Figure 7. Ligand density quality. Ligand density is shown with the published atomic model docked for deposited maps (left), and a re-built atomic model docked for our automated results (right), to emphasize the similar or improved map quality and interpretability.
Video 2. Ligand and water density in the binding site for the deposited maps/models (left) and automated maps/models (right). Note: no model was deposited with the apo map for EMPIAR-11800.

Looking to the ligand density for the deposited and automated maps in Figure 7, the density for most ligands is nearly indistinguishable or improved in interpretability in the automated maps. For example, automated processing for EMPIAR-11793 resolves a similar amount of the ligand and is overall less noisy. In addition to clearly defined ligand density, Video 2 shows the waters and ligand density in the binding pocket and compares the deposited to the automated processing maps. Density for water molecules in the ligand binding pocket are clearly resolved in both sets of maps.

Improvement in Model Building Metrics

To further confirm the aforementioned improvements in map quality, we remodeled the deposited atomic model against the automated density map for each of the datasets. With reduced anisotropy in the automated maps, we expected there to be less overfitting of the models due to anisotropy induced compression/stretching of the map, resulting in better metrics. After remodeling all datasets, including the apo-CAK dataset, we measured a series of map-model and map based metrics to quantify improvements.

Map-model metrics show that automated maps yield better models

Map-model Fourier Shell Correlation is a useful metric to determine how well a model agrees with a map. Figure 8 shows map-model FSCs for all map-model pairs (excluding apo-CAK, EMPIAR-11800; no model deposited). Map-model FSC curves are considered reliable up to a cutoff of 0.5. Comparing the curves for automated and deposited map-model pairs, there is clear improvement in medium to high resolution ranges for the automated maps compared to the deposited maps, as the FSC curves stay closer to 1.0. This indicates there is more agreement between the automated maps and models compared to the deposited maps/models.

Figure 8. Map-model FSCs for automated and deposited data showing the improvement in information content from DC to about 2.5Å. Deposited map-model FSCs (orange) and automated map-model FSCs (purple) computed from models refined into the automated maps. Map-model FSCs were computed using Phenix; gray, horizontal, dashed line at FSC = 0.5 for comparison.

We also quantified four other map-model parameters to see if other commonly tracked metrics improved. Using the validation tools in Phenix, we measured the CCmask, EMRinger, MolProbity, and Clash scores for all deposited and automated map-model pairs and plotted in Figure 9 (Liebschner et al., 2019).

Figure 9. Reduction in map anisotropy improves all measured map-model metrics. Phenix was used to compute CCmask (left), EMRinger (second column), MolProbity (third column), and Clash scores (right). Deposited map-model metrics are shown with gray circles while automated map-model metrics are shown with filled circles in green (improved), blue (same), and yellow (reduced). Note X-axis units for each plot as the axes are always plotted with improving values to the right.

CCmask scores measure the fit of atomic centers to the map, providing a quantitative understanding of model accuracy (Afonine et al., 2018). Using Phenix, we computed CCmask scores for the deposited and automated map-model pairs across all datasets. In every instance, our automated maps demonstrated improved model fits, with a maximum improvement of 0.16 observed for EMPIAR-11823d.

The EMRinger score evaluates the favorability of side-chain conformations relative to expected rotameric positions. It is quantified as the maximum normalized Z-score (adjusted for model length) across density thresholds, reflecting the agreement between the structural model and the map (Barad et al., 2015). Higher scores indicate superior model fit. We observed significant EMRinger score improvements ranging from 0.17 to 1.74 (averaging 0.77; n=8) for the automated map-model pairs compared to deposited results.

The MolProbity Score is computed using a “weighted function of clashes, Ramachandran favored, and rotamer outliers, scaled and normalized so that its value approximates the resolution at which that score would be average” (Williams et al., 2018). A higher score indicates more clashes, more Ramachandran outliers, and more rotamer outliers. All automated map model pairs have a lower MolProbity score with up to a 36% reduction for EMPIAR-11823c.

The Clash Score, also computed in Phenix, is derived from the MolProbity all-atom steric clash analysis and reported as the number of overlaps ≥0.4Å per thousand atoms (Afonine et al., 2018; Williams et al., 2018). Similar to MolProbity, lower values indicate fewer steric conflicts. For all automated datasets, we achieved a reduction in clashes by an average of 3 per thousand atoms; for models containing approximately 10,000 atoms, this corresponds to a total reduction of 30 clashes per model.

Per-residue cross-correlation improvements for the three most improved datasets

To further illustrate the map quality improvement on three of the most improved datasets, we evaluated the CCmask values on a more granular level by utilizing both the per-residue values and a 3-residue rolling window across all chains in the model (Figure 10): Chain H (Cyclin H), Chain I (MAT1), and Chain J (CDK7). We performed this analysis for EMPIAR entries 11823b, 11823d, and 11823ef. These values were calculated for every residue and water molecule modeled into the respective maps. Across these three datasets, it is evident that the CCmask values are consistently higher for the automated workflow results when compared to the manual processing outcomes. This represents a general increase across the entire map-model pair rather than better resolution in only a single chain or feature.

Figure 10. Cross correlation mask (CCmask) scores for three most improved datasets showing that the improvement is model wide and not due to a specific, local improvement. Deposited (orange) and automated (purple) CCmask scores on a 3-residue rolling average (dark lines) and per-residue data in translucent lines for each polypeptide (CDK7, MAT1, and Cyclin H) as well as all waters.

Automated curation isolates smaller, higher quality particle stacks

In most datasets, we observed that final particle counts from automated processing were lower than the deposited manual processing (Table 2). This is likely attributable to the rigorous automated curation steps required to eliminate map anisotropy. Our curation strategy prioritizes a strict acceptance of high-quality particles rather than maximizing the total particle count, ultimately yielding better results through more selective filtering.

Dataset
Deposited Final # Particles
Automated Final # Particles

11800

486,457

93,732

11793

555,117

468,128

11799

126,177

132,961

11807

176,991

167,148

11823a

294,085

88,447

11823b

334,554

97,538

11823c

294,740

103,098

11823d

565,148

82,165

11823ef

269,962

187,977

Table 2 - Particle counts used for the final reconstructions of deposited and automated maps. Particle counts for the automated workflow are smaller due to extensive curation and pruning of over represented views.

End-to-end processing times less than 24 hours on average

All processing for these datasets was completed on a single GPU node utilizing 32 CPU cores, 512 GB of RAM, and eight NVIDIA A100 (40GB) GPUs. We utilized parallelization across resources where possible; Reference Based Motion Correction (RBMC) which was restricted to 6 GPUs. The reported runtimes in Table 3 represent end-to-end processing for the entire workflow, inclusive of RBMC. For EMPIAR-11823e and 11823f, particle stacks were initially processed separately (requiring 16.4 and 16.6 hours, respectively) before being merged. We anticipate that total processing time for these combined datasets would align with 11793 or 11807 given the similar movie and particle counts. Average processing times across all datasets were approximately 21 hours.

Ultimately, for datasets collected in an unattended manner, by leveraging CryoSPARC Live for real-time preprocessing during collection and employing parallel processing strategies with automated workflows in CryoSPARC, it could be feasible to collect and process a fully loaded 12-grid autoloader within a handful of days.

Dataset
# of movies
Total processing time (hr)

11800

5,805

17.9

11793

9,118

23.9

11799

5,780

17.5

11807

10,017

25.2

11823a

5,729

18.7

11823b

5,583

18.3

11823c

5,210

17.0

11823d

5,203

18.0

11823ef

10,598

33.0

Table 3 - Number of movies and total processing time (hours) for all datasets interrogated in this study. These timings include all steps in the workflow.

arrow-up-left

How we iteratively developed an improved automation workflow to handle preferred orientation

Processing the apo-CAK dataset was undertaken with the explicit goals of understanding any data processing intricacies or pitfalls related to this target and generating a reference volume for further processing of drug bound CAK datasets. First, Workflow v1 was applied, but the jobs were not queued (Figure 18 - column 1). Each job in the workflow was manually run to process the apo-CAK dataset, allowing for inspection of results and to ensure parameters produce the desired outcomes before proceeding to the next job. Below, we explore these results along with observations made throughout processing that lead to an optimized workflow (Figure 18 - column 4) for automated processing of the other CAK datasets.

Initial workflow

Preprocessing

Micrograph preprocessing steps including Import Movies, Patch Motion Correction, Patch CTF Estimation, Manually Curate Exposures, Micrograph Junk Detector, and Micrograph Denoiser. Similar to the GPCR workflow, micrographs with a CTF fit resolution worse than 5Å and total full-frame motion greater than 200 pixels were excluded. This resulted in the rejection of 116 micrographs; leaving behind ~5,700 for the next processing steps. Micrographs put through this preprocessing pipeline result go from low contrast and difficult to interpret (Figure 11A) to easy to visualize particles and labeled junk (Figure 11B).

Figure 11A. Motion corrected micrograph of EMPIAR-11800 showing poor SNR for particles and a large swath of ethane contamination.
Figure 11B. The same micrograph as the left after junk is labeled and the micrograph is denoised and ready for particle picking.

Particle Picking

Having preprocessed micrographs, the Blob Picker was used with a minimum particle diameter of 60Å, a maximum diameter of 90Å, a minimum separation of 0.7 particle diameters to pick particles from 400 denoised micrographs. Following picking, particles were extracted with a box size of 300 pixels Fourier cropped to 128 pixels. Following blob picking, particle picks were filtered by first the Micrograph Junk Detector to remove particle picks on junk and then by the auto mode of Inspect Particle Picks. The power score for auto Inspect Particle Picks had to be reduced to 20 due to the smaller size of the particle and lack of a micelle compared to GPCRs. Particles were then subject to 2D Classification to generate templates for template picking.

Figure 12. Poor orientation distribution in 2D-classification utilizing blob picking. 2D classes of particles picked from 400 micrographs.

Right away there were two things to note about some classes in Figure 12:

  1. Classes that were higher resolution and had clear secondary structural features represented a very similar viewing direction of the CAK complex.

  2. A few classes at the bottom indicated there were hot pixels present. After some investigation we found that these came from a defective gain file. Outliers in the gain file were set as defective pixels and movies were re-motion corrected before continuing with processing.

Processing was continued at this point where good templates, Figure 13, were selected from the above 2D classification to be used for template picking. Although we recognized this would likely lead to a poor reconstruction, this was done to demonstrate what would happen if poor templates were used to pick particles. These templates were used to template pick all micrographs in the dataset. Once all micrographs were picked, particle picks were put through the Micrograph Junk Detector and the automated modes of Inspect Particle Picks to remove picks on junk or on the background resulting in a particle stack that was roughly 2.1M particles. These particles were extracted from the micrographs in a box size of 384 pixels and Fourier cropped to 128 pixels.

Figure 13. Orientation bias is introduced when selecting templates from the above 2D-classification job. Classes selected from 2D Classification of blob picked particles. Notable, only a few views, mostly of the broad side of the CAK complex.

Particle Curation

Since we were processing the first dataset in the series (apo-CAK), we did not have a reference volume to use in a decoy classification. Therefore, the cleaned particle stack was first subjected to multi-class Ab-Initio Reconstruction to generate a set of volumes to use for heterogeneous refinement. For this Ab-Initio Reconstruction, the initial resolution and maximum resolution are increased to 20Å and 5Å, respectively. These resolution choices were inspired by a recent preprint about High-resolution Ab-Initio Reconstruction (HR-HAIR; Kim et al., 2025) and our development of Homogeneous Ab-Initio Refinement in CryoSPARC v5 . We found that to get a successful volume that would lead to desirable results from Heterogeneous Refinement, the 5Å maximum resolution was key. Even decreasing the maximum resolution to 8Å resulted in unusable volumes. Full parameter selections for Ab-Initio Reconstruction are below, and these were used for all Ab-Initio Reconstruction jobs moving forward unless otherwise noted.

Parameter
Setting
Explanation

Number of Ab-Initio classes

3

3 classes seemed appropriate since the sample was relatively pure and there were few components present

Number of particles to use

250,000

Using 10% of the initial particle stack usually results in a decent set of volumes

Maximum resolution

5

Smaller particles typically benefit from higher resolutions - generating better volumes

Initial Resolution

20

Following multi-class Ab-Initio Reconstruction, all particles and volumes were subjected to Heterogeneous Refinement with default parameters. Figure 14 shows one class (green) depicting CAK that was manually selected to move forward with in processing, but it is highly anisotropic.

Figure 14. Initial volumes after the first Heterogeneous Refinement, displayed at the same threshold standard deviation. The green volume appears to be the shape of CAK, but is extremely anisotropic.

Following this, duplicates were pruned from the particle stack of the selected green class, and then were re-extracted at a box size of 384 pixels and Fourier cropped to 256 pixels. These particles were then subjected to a second multi-class Ab-Initio Reconstruction using the following parameters:

Parameter
Setting
Explanation

Number of Ab-Initio classes

5

Since there is significant anisotropy, opted for more classes

Number of particles to use

100,000

Using ~10% of the initial particle stack which usually results in a decent set of volumes

Maximum resolution

5

Smaller particles typically benefit from higher resolutions - generating better volumes

Initial Resolution

10

Inspired by the HR-HAIR preprint, we opted to begin perform pose estimation at medium resolutions

All particles and volumes from the Ab-Initio Reconstruction were used for another Heterogeneous Refinement using the following settings:

Parameter
Setting
Explanation

Refinement box size

256

Using a larger box size allows higher resolution reconstructions

Batch size per class

3000

Increasing this number can provide better classification results and this value was used in the GPCR workflow

Initial Resolution

12

Lower resolutions lead to poorer results in initial iterations

The best class (Figure 15: green) was selected for further processing. Since there was only one good class, the particles associated with this class were reconstructed with a Homogeneous Reconstruction job and refined using Non-uniform Refinement. After Non-uniform Refinement the map exhibited anisotropy, and the density in many regions was discontinuous (Figure 19, column 1). Additionally, the low cFAR score (0.09) matched the quality of the reconstruction.

Figure 15. Volumes from second round of particle curation (heterogeneous refinement), with the green being the good volume that moved through the processing pipeline. Green volume was selected for further processing. Other volumes present, notable class 0 and class 3 (left to right) look like CAK, but are highly anisotropic.

Addressing preferred orientation

At this point, all jobs related to particle curation in the Workflow v1 had been successfully run, but a low cFAR score indicated there was remaining map anisotropy that needed to be addressed. First, the Rebalance Orientations job type was used to remove particles from viewing direction bins that were over-represented. The Non-uniform Refinement leading into the Rebalance Orientations job showed a slight shoulder of low scale particles emerging and thus overrepresented particles with the lowest per-particle scale (PPS) values were removed from their viewing direction bin. Although it appears that many particles were removed with scales greater than 1.0, it should be noted that the particles being removed are those that are overrepresented and thus would be considered “good particles”, but in comparison to their binned counterparts, have a lower per-particle scale. Since the resulting rebalanced particle stack has a clear, bi-modal distribution (Figure 16C,D), Subset Particles by Statistic was used to to remove low scale particles, and retaining particles in Cluster 1 (light blue) for further processing.

Figure 16. Rebalanced orientations and per-particle scale values for green map of Figure 14 showing over-represented views that can be pruned along with the low-scale particles. (A,B) Rebalancing the orientations reduced the number of over represented views present in the dataset. (C) After rebalancing orientations, a bimodal distribution of per-particle scale values are observed (blue) once the over represented views (red) are removed. (D) Bi-modal distribution of CAK per-particle scale values, clustered into low scale (dark blue) and high scale (light blue) by the job Subset Particles by Statistic.

The resulting particle stack was then subjected to a final Homogeneous Reconstruction and Non-Uniform Refinement. For the Non-Uniform Refinement, global (3rd order only) and local CTF corrections were enabled. The resulting volume had less visible anisotropy and the cFAR score improved from 0.09 to 0.31 (Figure 19, column 2). Although this map looked better than the first, further approaches were undertaken with changes to particle picking and curation steps with hopes to reduce anisotropy and improve the connectivity of map density.

Intermediate workflow changes

Still processing in a semi-manual manner (selecting volumes manually to move through the workflow), we aimed to improve the resulting map anisotropy. We hypothesized that changing the picking strategy might pick more rare views, especially since the templates we used in the first workflow were all of one view of the CAK complex. Therefore, using the Blob Picker with the following parameters, picking was carried out on all micrographs using circular blobs.

Parameter
Setting
Explanation

Minimum particle diameter (Å)

40

Based on the previous reconstruction, this could be the smallest view

Maximum particle diameter (Å)

90

Based on the previous reconstruction, this could be the largest view

Pick on denoised micrographs

True

Picking on denoised micrographs yields better picks and allows filtering via auto-Inspect Particle Picks

Min. separation dist (diameters)

0.6

Likely too low, but the results of the picking job looked good upon manual inspection

After particle picking, picks were cleaned as before using the Micrograph Junk Detector and Auto-Inspect Particle Picks. Next, a similar set of particle curation steps were utilized with a few changes. The first change was using a funnel-like approach to filtering poor particles where the first multi-class Ab-Initio Reconstruction and Heterogeneous Refinement used five classes and the second round used three classes. Between the two rounds of 3D particle curation, another Rebalance Orientations job was included to prune out overrepresented views early on. This helped to reduce anisotropy before the second round of particle classification.

Looking to the effects of these above changes, Figure 19 (column 3) shows that the effects of these changes were minimal, but in the correct direction. Resolution and cFAR scores improved but the quality of density was roughly the same.

Final Workflow Changes: Workflow v2

As a final approach to reducing anisotropy and improving the density of the final reconstruction, we opted for a different particle picking strategy and initial curation strategy. Both of these improvements utilized the apo-CAK volume that was output at the end of intermediate workflow. This volume was lowpass filtered to 15Å to reduce any high resolution bias for the two major changes to the workflow.

Using the apo-CAK volume output at the end of the intermediate workflow, the Create Templates job with default parameters was used to create 50 templates, uniformly spaced on the viewing sphere, to use for particle picking (Figure 17).

Figure 17. Templates created from LPF apo-CAK volume output from the second workflow.

These templates were used with the Template Picker on denoised micrographs and the below parameters to fully pick all micrographs. After particle picking, picks were cleaned up using the Micrograph Junk Detector and Auto-Inspect Particle Picks.

Parameter
Setting
Explanation

Particle diameter (Å)

90

Based on the previous reconstruction, this could be the largest view

Pick on denoised micrographs

True

Picking on denoised micrographs yields better picks and allows filtering via auto-Inspect Particle Picks

Use CTFs to filter templates

False

Automatically disabled when picking on denoised micrographs

Min. separation dist (diameters)

0.6

Provides a minimum separation distance of 54Å

To help ensure that rare views were correctly classified into a good class during particle curation, we changed the first round of particle curation from Ab-Initio Reconstruction and Heterogeneous Refinement to “decoy classification”. Decoy classification is the use of known volumes (one representing the target class plus a number of junk volumes to sequester poor particles) as inputs into a Heterogeneous Refinement job (see here for details). Decoy classification is useful for (a) providing a reference accurate enough to capture all relevant views of the particle and (b) easily enabling automated selection of good particles (as they presumably end up in the good class). The low resolution reference of apo-CAK lowpass filtered to 15Å was used as the good volume and the same junk volumes from the GPCR datasets (see Table 1 of the preprint) were used for decoy classification with the apo-CAK particle stack.

Following decoy classification, particles in the good class were globally refined to obtain more accurate pose information and to compute per-particle scale values before pruning over-represented views with Rebalance Orientations. As is the same in intermediate workflow, particles were then subjected to a second round of curation using Ab-Initio Reconstruction and Heterogeneous Refinement with the same parameter settings as previous workflows. Use of Reference Based Auto Select 3D after the second round of curation allowed for automatic selection of the good class(es) based on resolution. Particles from the high-resolution, good classes were then reconstructed, globally refined, and particles from over-represented views and particles with low per-particle scale values were removed before performing Reference Based Motion Correction and a final Non-uniform Refinement.

After implementing these changes, the workflow was run against the apo-CAK dataset. The resulting volume had more connectivity, achieved higher resolution, and the cFAR score was further improved from the intermediate workflow, 0.35 to 0.40 (Figure 18 - Final Workflow, Figure 18 - column 4).

Workflow evolution

Figure 18 details all incremental changes made to Workflow v1, which ultimately yielded Workflow v2. In totality, the particle picking stage went through the most changes and the particle curation stages feature the inclusion of two Rebalance Orientations job and one Subset Particles by Statistic.

Figure 18. Workflow evolution across trials - (Column 1) Workflow v1 applied to CAK. (Column 2) Small addition to Workflow v1 to reduce anisotropy encountered in the CAK datasets. (Column 3) Second workflow trying to improve map anisotropy. (Column 4) Final workflow settled upon for automated data processing of CAK (Workflow v2). Jobs with black outlines constitute changes from the previous column.

The changes to the workflow can be directly visualized in Figure 19, wherein the map anisotropy is reduced at each version of the workflow as judged by the cFAR scores. In addition, the map density displays more connectivity in the final volume of the workflow in comparison to the first.

Figure 19. Clear improvements in the map quality and cFAR score between workflows. GSFSC curves and cFSCs for all four volumes were computed with the same mask for fair comparison.

Conclusion

This case study demonstrates that automation in CryoSPARC is robust and generalizable, and can produce results that are equal or better than manual processing, even on a challenging series of datasets of a small drug-bound particle with strong preferred orientation. The updated Workflow v2 used here can be downloaded and easily used to automate processing of your own targets.

Automation of data processing in CryoSPARC enables significant reductions in time-to-structure, particularly for SBDD scenarios. Processing times on the order of 20 hours can be obtained on modest, local compute infrastructure. Furthermore, for a scientist to adapt our published workflows to their own target, the time investment is minimal, typically requiring only one or two days of exploratory processing with an exemplar dataset. After this initial work, all subsequent datasets of the same target can be processed fully automatically.

How to set up automation in CryoSPARC for a new target class

All of the methods and tools required to set up an automated workflow for a new target class, such as CAK, are already available in CryoSPARC v4.7.1+. A workflow file, reference volume, final maps, and final models from the automation study are available here:

To apply the Workflow to a new target, you will need a reference volume of your target from a previously processed dataset. Granted these items are available, you can:

  1. Download the workflow JSON files and junk volumes: automated_workflow_materials_v2.zip

  2. Upload these volumes to your compute setup.

  3. Import workflow JSON file into CryoSPARC using the workflow panel in the sidebar.

  4. Select the workflow from the list and set all flagged parameters. Parameters that might need to be adjusted (on a per dataset basis) are below:

    Import movies

    • Movies data path

    • Gain reference path

    • Raw pixel size (Å)

    • Accelerating voltage (kV)

    • Spherical aberration (mm)

    • Total exposure dose (e/Å^2)

    • Flip gain ref and defect file in Y

    • Exposure group metadata (if present)

    Import 3D Volumes

    • Path to reference and junk volumes

    Patch Motion Correction

    • Output F-crop factor

    Template Picker

    • Particle diameter (Å) - set to the largest particle dimension

    Extract from micrographs (2x)

    • Extraction box size - set to a value that is 2-3x the size of your target

  5. Click on the green “Apply” button in the bottom of the workflow GUI.

  6. Queue and run each job independently, or in small clusters. Review all results to ensure that desired outputs are generated and there are no processing issues. A few key results to check include:

    1. Template Picker - ensure minimum separation distance setting resulted in adequate picking (not to sparse, not double picking particles)

    2. Inspect Particle Picks - ensure that the power value selected is correct for the peak of particles (see GPCR automation preprint for more info)

    3. Rebalance Orientations - on a target that exhibits little or no preferred orientation, Rebalance Percentile can be increased to a value of 80-90 so as to preserve more of the picked particles.

    4. Extract from Micrographs (2nd instance) - modulate the Fourier crop box size parameter to ensure that the resolution of reconstruction is not inhibited by Nyquist (more on this here).

  7. If issues arise, add exploratory processing steps to troubleshoot these issues.

  8. Once all processing is completed to a satisfactory result, select all jobs, rebuild the workflow, and apply it to other datasets of the same target.

If you collected data using beam image-shift, include an Exposure Group Utilities job between Manually Curate Exposures and the Micrograph Denoiser.

References

Cushing, V. I. et al. High-resolution cryo-EM of the human CDK-activating kinase for structure-based drug design. Nat Commun 15, 2265 (2024).

Liebschner, D. et al. Macromolecular structure determination using X-rays, neutrons and electrons: recent developments in Phenix. Acta Crystallogr D Struct Biol 75, 861–877 (2019)

Afonine, P. V. et al. New tools for the analysis and validation of cryo-EM maps and atomic models. Acta Crystallogr D Struct Biol 74, 814–840 (2018).

Barad, B. A. et al. EMRinger: side chain–directed model and map validation for 3D cryo-electron microscopy. Nat Methods 12, 943–946 (2015).

Williams, C. J. et al. MolProbity: More and better reference data for improved all‐atom structure validation. Protein Science 27, 293–315 (2018).

Kim, K., Li, H. & Clarke, O. B. High-resolution ab initio reconstruction enables cryo-EM structure determination of small particles. bioRxiv (2025).

Last updated