> For the complete documentation index, see [llms.txt](https://guide.cryosparc.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://guide.cryosparc.com/processing-data/automated-workflows/advanced-automated-data-processing-a-case-study-using-cak.md).

# Advanced automated data processing: a case study using CAK

## Introduction

[Automated cryo-EM data processing is an important area of development](https://guide.cryosparc.com/processing-data/automated-workflows#automated-repeat-target-structure-determination) because it can reduce time-to-structure, increase the number of parallel datasets that can be processed, lower the barrier to entry for new practitioners and ensure reproducible results across experiments.

We initially [demonstrated an automation strategy with a test set of 21 challenging GPCR datasets](https://www.biorxiv.org/content/10.1101/2025.10.17.682689v1), processed fully automatically, using a general automation Workflow and strategy that can be easily adapted to new target classes (”General Automation Workflow v1”). The Workflow and tools that were developed in CryoSPARC to support automated data processing were able to produce results **equal or better than manual processing** with **zero manual intervention** for repeat target scenarios.

To further demonstrate generalizable automation, and to serve as a guide for users applying CryoSPARC automation to their own targets, **this case study covers the application of automation to a series of nine drug-bound CDK-activating Kinase (CAK) datasets**. CAK is a very small, soluble, and druggable protein complex that is an oncogenic and antiviral target of interest due to its central role in physiology. Our efforts to automate data processing build upon the original CAK dataset depositors' ([Cushing et al., 2024](https://www.nature.com/articles/s41467-024-46375-9)) goal of establishing high-throughput automated data collection, and together these improvements can help accelerate drug discovery workflows, especially for small targets like CAK.

**We provide an updated and improved version of a general automation Workflow, “General Automation Workflow v2”, and show that it produces equal or better results on CAK datasets compared to manual processing.** The updated Workflow is more robust to preferred orientation than the previous version and can be [downloaded and used to automate processing of your own target](#how-to-set-up-automation-in-cryosparc-for-a-new-target-class).

The case study is organized as follows:

* [**Background about CAK**](#background-on-cdk-activating-kinase-cak-and-description-of-test-datasets) and details about the 9 CAK datasets, including the severe preferred orientation challenges they present
* [**An updated automation workflow (Workflow v2)**](#an-updated-automation-workflow-workflow-v2) simplified and improved from the workflow we initially used for GPCRs (Workflow v1) and able to handle severe preferred orientation
* [**Results from using the updated workflow on CAK**](#fully-automated-processing-results-for-cak) datasets, demonstrating that we can achieve fully automated results, where the map quality exceeds that of manual, expert processing
* [**A detailed step-by-step explanation of how we developed the updated workflow**](#how-we-iteratively-developed-an-improved-automation-workflow-to-handle-preferred-orientation), iteratively by processing an apo-CAK dataset and adding stages to handle preferred orientation
* [**A simple walkthrough of how to set up automation**](#how-to-set-up-automation-in-cryosparc-for-a-new-target-class) in CryoSPARC for your own target class

{% hint style="success" icon="check" %}
All of the jobs, methods, and tools required to use and set up automated workflows, including the workflow presented in this case study, are available in CryoSPARC v4.7.1+. [**You can download the Workflow v2 JSON file**](#how-to-set-up-automation-in-cryosparc-for-a-new-target-class)**, and import it into your own CryoSPARC instance to begin automated processing of your own data.**
{% endhint %}

## Background on CDK-activating Kinase (CAK) and Description of Test Datasets

CAK (CDK-Activating Kinase) is a critical regulator of cell division and growth. This complex consists of three subunits—MAT1, Cyclin H, and CDK7—wherein ligands bind specifically to a cleft in CDK7 (**Figure 1**, green ligand). Despite being a significant target for antiviral and cancer therapies, CAK remains a notoriously challenging cryo-EM subject due to its small size (approximately 80 kDa).

<figure><img src="/files/5ylliWMejhD14AXX4Mpm" alt="" width="375"><figcaption><p>Figure 1. Cartoon of the CAK complex highlighting many important features. Low resolution surface representations of the CAK complex including CDK7 (blue), Cyclin H (grey), MAT1 (orange), and the ligand (green). All ligands bind in the same cleft as the ATPγS molecule depicted.</p></figcaption></figure>

This case study uses nine EMPIAR datasets, encompassing CAK in its apo state, bound to a nucleotide analog, and in complex with seven distinct inhibitors (Table 1). The datasets featured in this case study were originally processed by the depositors using a manual approach involving CryoSPARC (versions 3.3.1-4.1.1) and RELION 4.0. To give a brief overview of the depositors data processing approach, micrograph preprocessing were completed during data collection within [CryoSPARC Live](https://guide.cryosparc.com/live/about-cryosparc-live). Particles were then picked using the Blob Picker and subjected to 2D Classification; ultimately, particles from the selected classes were reconstructed in CryoSPARC Live using Homogeneous Refinement to yield a starting volume. Once data collection was finalized, the depositors utilized templates from the CryoSPARC Live Streaming 2D Classification, as well as templates generated from the volume output of the CryoSPARC Live Homogeneous Refinement, to perform template picking across the entire dataset in separate jobs. All blob and template-picked particles underwent duplicate removal before being converted to the STAR file format and exported to RELION. Subsequently, a series of 3D auto-refinements, 3D classifications, global CTF refinements, and Bayesian particle polishing jobs were executed to achieve the high-resolution reconstruction utilized for model building. While this represents the general workflow employed by the depositors, they noted that certain targets and samples necessitated additional, specialized steps.

<table><thead><tr><th width="90.203125">EMPIAR</th><th width="80.34765625">EMDB</th><th width="79.7734375">PDB</th><th width="131.640625">Inhibitor</th><th width="130.3515625">Grid name</th><th width="96.5703125"># of movies</th><th width="168.890625">Final # of particles yielding deposited map</th></tr></thead><tbody><tr><td>11800</td><td>17523</td><td>---</td><td>Apo</td><td>VC15-1</td><td>5,805</td><td>486K</td></tr><tr><td>11793</td><td>17129</td><td>8ORM</td><td>THZ1</td><td>29-1</td><td>9,118</td><td>555K</td></tr><tr><td>11799</td><td>17511</td><td>8P6Y</td><td>ATPgS</td><td>VC16-3</td><td>5,708</td><td>126k</td></tr><tr><td>11807</td><td>17754</td><td>8PLZ</td><td>CT7030</td><td>BG51-1</td><td>10,017</td><td>177K</td></tr><tr><td>11823a</td><td>17520</td><td>8P77</td><td>ICE-C0943</td><td>VC1-1</td><td>5,729</td><td>294k</td></tr><tr><td>11823b</td><td>17515</td><td>8P72</td><td>ICE-C0768</td><td>VC2-4</td><td>5,583</td><td>335K</td></tr><tr><td>11823c</td><td>17521</td><td>8P78</td><td>Dinaciclib</td><td>VC4-3</td><td>5,210</td><td>295K</td></tr><tr><td>11823d</td><td>17509</td><td>8P6W</td><td>BS181</td><td>VC8-1</td><td>5,203</td><td>565K</td></tr><tr><td>11823ef</td><td>17508</td><td>8P6V</td><td>ICE-C0942(a)</td><td>VC13-1/VC14-1</td><td>10,598</td><td>270K</td></tr></tbody></table>

<sup>**Table 1.**</sup> <sup>**CAK Automation Dataset Information**</sup> <sup></sup><sup>- Information for all nine datasets used in this study including the EMPIAR, EMDB, and PDB depositions. All datasets were collected on a Krios G4 operating at 300 keV and equipped with a Falcon 4i and SelectrisX energy filter. All EER formatted movies were collected using the same data collection parameters: pixel size of 0.57Å, a dose of 70e-/Å2, 70 fractions, and on a microscope operating Cs value of 2.7, and at 300 keV. Dataset size \~2.2 TB/5,000 movies.</sup>

The CAK datasets described above are challenging to process for a few reasons. The CAK complex is small (\~85 kDa) and therefore the signal-to-noise (SNR) ratio, which affects both particle picking and alignment, is limited and drops quickly in thicker ice. While the depositing authors were able to make grids with very thin ice and therefore collect datasets with improved SNR, the grids presented significant preferred orientation which hindered the final quality of the reconstructions. Looking below to **Figure 2**, although we can see that while high resolution features are present, including holes in some aromatic rings, there is stretching and fragmentation of the density across the map in the direction of the poorly resolved orientation. Global orientation metrics for the deposited map, such as the low cFAR score and the missing wedge of information in the 3DFSC central slice, highlight the challenges presented by these datasets. Many cryo-EM datasets have similar issues, and preferred orientation is a common reason for failure of cryo-EM reconstruction.

<figure><img src="/files/vs4iNpRmafMReAXfdXCm" alt=""><figcaption><p><strong>Figure 2. Anisotropy present in the deposited CAK maps presents a challenging case to deal with preferred orientation in an automated manner.</strong> (A) Face and side views of the deposited map for EMPIAR-11823d highlighting the high resolution features, but also the presence of the anisotropic stretching. (B) cFAR plot for the deposited map. (C) Plot of all 3072 cFSCs showing the wide range of resolutions across the viewing sphere. (D) Central slices of 3DFSC volume showing the missing wedge of information due to preferred orientation.</p></figcaption></figure>

## An Updated Automation Workflow: Workflow v2

The general automation strategy that we use for fully automated data processing ([described in detail in this preprint](https://www.biorxiv.org/content/10.1101/2025.10.17.682689v1)) involves standard preprocessing, high-quality generalizable particle picking, extensive and robust particle curation, and multiple refinements to achieve the highest quality results. In our initial demonstration of this strategy for repeat GPCR structure determination, we developed a CryoSPARC Workflow ([Automation Workflow v1](https://guide.cryosparc.com/processing-data/automated-workflows/preprint-end-to-end-fully-automated-data-processing-of-gpcrs)) that could be easily generalized to different target classes by simply providing a new target-class-specific low-resolution (15Å) reference volume.

For CAK, we used [Workflow v1](https://guide.cryosparc.com/processing-data/automated-workflows/preprint-end-to-end-fully-automated-data-processing-of-gpcrs) as a starting point for processing a single apo-CAK dataset ([EMPIAR-11800](https://www.ebi.ac.uk/empiar/EMPIAR-11800/)), and found that it produced results similar to the deposited maps, but ran into the challenges outlined above, in particular, strong preferred orientation present in the data. We therefore modified and updated the automation workflow, following the same overall strategy but simplifying particle picking and adding additional steps to account for preferred orientation. The experimental development leading to this updated workflow, Workflow v2, [is detailed in a section below](#how-we-iteratively-developed-an-improved-automation-workflow-to-handle-preferred-orientation).

Workflow v2 (**Figure 3**) provides simplified, robust, and fully automated end-to-end processing. Thanks to the [general automation strategy](https://www.biorxiv.org/content/10.1101/2025.10.17.682689v1.full) on which it is based, Workflow v2 can be **run automatically on new datasets with no changes** (other than dataset-level inputs such as microscope parameters). It can also be **generalized to new target classes by changing only the target-class/workflow level inputs** (a low resolution reference volume, particle diameters, thresholds, and junk volumes).

<figure><img src="/files/CRZ75420dfsJTGupXqu7" alt=""><figcaption><p><strong>Figure 3. Fully-automated workflow (Workflow v2) developed for processing all CAK datasets that relies on modified picking and curation strategies to reduce effects of preferred orientation.</strong> This figure shows a majority of jobs in the workflow along with all required inputs for those jobs, and it can be easily adapted to new targets with a few key adjustments to the highlighted parameters.</p></figcaption></figure>

As shown in **Figure 3**, Workflow v2 contains all end-to-end data processing steps. Preprocessing steps including [Patch Motion Correction](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction) and [Patch CTF Estimation](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/ctf-estimation/job-patch-ctf-estimation#implementation-details) followed by the [Micrograph Denoiser](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/exposure-curation/job-micrograph-denoiser-beta) and [Micrograph Junk Detector](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/exposure-curation/job-micrograph-junk-detector-beta#particles-accepted). Using the [Template Picker](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/particle-picking/job-template-picker#description), particles are picked using the templates created from the input low resolution reference volume and a user supplied particle diameter and separation distance. Particle picks are then cleaned using the auto option in [Inspect Particle Picks](https://guide.cryosparc.com/application-guide/interactive-jobs#:~:text=the%20exposure%20viewer.-,Interactive%20Job%3A%20Inspect%20Particle%20Picks,-The%20inspect%20particle) to remove particle picks on the background and picks on junk regions are removed using the Micrograph Junk Detector.

After particle picks have been cleaned, particles are extracted and submitted to decoy classification — [Heterogeneous Refinement](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/3d-refinement/job-heterogeneous-refinement#description) using as inputs the cleaned particle stack, the low resolution reference volume, and four junk volumes to attract poor particles. Note that the junk volumes we use for CAK are identical to those used for GPCR datasets in Workflow v1. The good volume and particle stack are then subjected to a [Non-uniform Refinement](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/3d-refinement/job-non-uniform-refinement-new#common-parameters) to improve particle poses and compute per-particle scale values. [Rebalance Orientations](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/particle-curation/job-rebalance-orientations#rebalance-percentile) is then used to filter out particles in over-represented viewing directions. A second round of particle curation is then performed on the rebalanced particle stack where a three-class [Ab-Initio Reconstruction](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/3d-reconstruction/job-ab-initio-reconstruction#common-problems) is used followed by Heterogeneous Refinement. Using [Reference Based Auto Select 3D](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/variability/job-reference-based-auto-select-3d-beta#reference-volume), the highest resolution volume(s) and particle stack(s) are selected for additional pruning of over-represented views using Rebalance Orientations followed by removal of low scale particles with [Subset Particles by Statistic](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/particle-curation/job-subset-particles-by-statistic#subset-by). Together, these additional stages greatly reduce map anisotropy by pruning over-represented views and removing low-quality particles.

{% hint style="info" icon="memo-pad" %}
**Note:** Workflow v2 makes use of Rebalance Orientations to prune heavily overrepresented views. If you are using Workflow v2 on a target that exhibits little or no preferred orientation, the Rebalance percentile can be increased to a value of 80-90 so as to preserve more of the picked particles.
{% endhint %}

For final refinements, particles are subjected to Non-uniform Refinement, [Reference Based Motion Correction](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/motion-correction/job-reference-based-motion-correction-beta#references), and a final Non-uniform Refinement. Here, [on-the-fly global CTF refinement in the Non-uniform Refinement job](https://guide.cryosparc.com/processing-data/tutorials-and-case-studies/tutorial-ctf-refinement#on-the-fly-ctf-refinement-in-homogeneous-refinement:~:text=alignments%20of%20particles.-,Non%2Duniform%20refinement%20with%20high%2Dorder%20CTF%20correction,-In%20CryoSPARC%20v2.12) is used to correct for residual beam tilt left over from AFIS data collection. Figure 4, below, details the differences between the v1 and v2 versions of the workflow.

<figure><img src="/files/zRQ4ZuJ7YMQUbjJTiACV" alt="" width="375"><figcaption><p><strong>Figure 4. Changes made between Workflow v1 and Workflow v2.</strong> Workflow v1 was originally developed for general automation and applied to a test set of 21 GPCRs and Workflow v2 is an improved iteration applied to the CAK datasets in this case study. The overall structure of the Workflows is similar; the picking strategy was changed in Workflow v2 and we introduced Rebalance Orientations and Subset Particles by Statistic to increase robustness to severe preferred orientation.</p></figcaption></figure>

In the next section, we show results of automated processing of CAK datasets. To confirm the generalizability of Workflow v2, we also used it to reprocess GPCR datasets and found that it produces similar or improved results compared to the [original Workflow (v1)](https://guide.cryosparc.com/processing-data/automated-workflows/preprint-end-to-end-fully-automated-data-processing-of-gpcrs).

## Fully Automated Processing Results for CAK

Using Workflow v2, we processed all eight ligand-bound CAK datasets along with the apo dataset in a fully hands-off manner. Below we summarize map metrics and improvements to map quality and the resulting models that stem from a higher quality reconstruction.

#### **Data and Input Preparation**

We processed the apo-CAK dataset manually by following the stages of Workflow v1 [(see below)](#initial-workflow) and the resulting refined volume, low-pass filtered to 15Å, became the reference volume used for the other eight CAK datasets. The reference volume is used in two places in the automated workflow (**Figure 3**), to create templates for template picking, and as the good volume during decoy classification.

Automated processing requires setting dataset-level, target-class-level and workflow-level inputs. Dataset names used below are based on their EMPIAR ID, and for EMPIAR-11823, the names are assigned alphabetically (a-f) based on the order of the datasets in the deposition. Full dataset information can be found in **Table 1**. Dataset-specific microscope parameters for each dataset were obtained from the respective EMPIAR depositions and these values were the same for all datasets: 0.57 Å pixel size, 70 electrons/Å<sup>2</sup>, and EER format fractionated into 70 frames. Next, based on processing the apo-CAK dataset, multiple class and workflow level parameters were selected. For particle picking, a particle diameter of 90Å and a separation distance of 0.6 diameters (54Å) were selected. For initially cleaning particle picks, it was necessary to adjust the power score of the Inspect Particle Picks auto-cluster mode to 20 since the CAK complex is soluble (lacks a micelle) and is much smaller than a GPCR. The final refinement map of apo-CAK was lowpass filtered to 15Å and used in conjunction with the GPCR junk volumes for decoy classification. Unlike processing the GPCR datasets where the class-level parameters required changes, these target datasets are all extremely similar and thus none of these parameters were varied between them.

#### **FSC Resolution and Map Anisotropy**

To be useful in practice, automated data processing should produce results equal or better than manual processing in terms of map quality and downstream interpretability. The first measures of success of automated data processing of the CAK complex datasets were the resolution and cFAR scores of the final volumes; summarized below in **Figure 5**. Comparing deposited manual processing results to our automated results, the resolution of the apo dataset was improved by 0.1Å, was unchanged for EMPIAR 11799, and was within 0.1Å to 0.3Å for all other datasets. Since global resolutions do not always provide a complete picture for datasets that have preferred orientation, we also examined map anisotropy. [cFAR scores](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/utilities/job-orientation-diagnostics#cfar-and-scf-at-a-glance:~:text=find%20missing%20views), show a major improvement in the automated results: values are near-zero for most deposited maps, indicating poor map quality due to preferred orientation, versus an average increase of 0.33 for automated maps, indicating interpretable and higher quality map density.

<figure><img src="/files/hQcTWOIgH4AiKjSep7lJ" alt=""><figcaption><p><strong>Figure 5. Global map metrics confirm that automated maps have similar resolution but significantly improved orientation distributions compared to deposited maps.</strong> Left: Gold-standard Fourier Shell Correlation (FSC) and right: Conical FSC Area Ratio (cFAR). Values from deposited published results shown as gray circles. Values from automated processing shown as filled, colored circles: improved (green), equivalent (blue), or reduced (orange).</p></figcaption></figure>

#### **Improved Map Quality and Interpretability**

To further evaluate the results of automated processing, we examined visual map quality and how well the ligand was resolved. In **Figure 6**, we have summarized the results of the deposited, manually processed data alongside our results from automated data processing. Map pairs are shown at the same number of standard deviations from the mean (9) for equivalent comparison. All maps, both deposited and automated, contain sidechain density, and for some maps, holes in large aromatic residues are present. In all cases, all three protein chains for CDK7, Cyclin H, and MAT1 are well resolved, and the ligand is present in the density.

For most of the deposited maps, the face views do not appear to exhibit anisotropy, but in the side views, there is strong anisotropy that manifests as stretching of the density along with increased discontinuity in regions, especially density for the protein backbone. **Video 1** shows this phenomenon for EMPIAR-11823d where upon rotation of the volume, the comparison is quite striking. Comparing the deposited maps to the automated maps, the map anisotropy is greatly reduced in the automated results and these observations of anisotropy are confirmed by the improvement in the cFAR scores in **Figure 5**.

<figure><img src="/files/5rcdcKgK75OPDimctegV" alt=""><figcaption><p><strong>Figure 6. Automated maps from the final Non-uniform Refinement exhibit well resolved density and less anisotropic effects compared to the deposited maps.</strong> Deposited maps and our automated processing maps are shown at similar thresholds for all datasets. Map coloring follows the same color scheme as Figure 1. For each global map, FSC resolutions (left, boxed) and cFAR score (left, unboxed) are shown. For automated maps, color scheme denotes improved (green), equivalent (blue), or worse (orange) relative to the manually processed deposited maps.</p></figcaption></figure>

<figure><img src="/files/8PAywlWwj5wmqAuJ3j5i" alt=""><figcaption><p><strong>Video 1. Map density from automated processing is of higher quality as compared with the deposited density.</strong> This video shows the reduction of the map anisotropy for the automated processing of EMPIAR-11823d (left) compared to the deposited map (right).</p></figcaption></figure>

<figure><img src="/files/DlhpNxwrQTcKJwOJSFLg" alt="" width="375"><figcaption><p><strong>Figure 7. Ligand density quality</strong>. Ligand density is shown with the published atomic model docked for deposited maps (left), and a re-built atomic model docked for our automated results (right), to emphasize the similar or improved map quality and interpretability.</p></figcaption></figure>

<figure><img src="/files/neAVegXo3hwDdQDRruak" alt=""><figcaption><p><strong>Video 2. Ligand and water density in the binding site for the deposited maps/models (left) and automated maps/models (right).</strong> Note: no model was deposited with the apo map for EMPIAR-11800.</p></figcaption></figure>

Looking to the ligand density for the deposited and automated maps in **Figure 7**, the density for most ligands is nearly indistinguishable or improved in interpretability in the automated maps. For example, automated processing for EMPIAR-11793 resolves a similar amount of the ligand and is overall less noisy. In addition to clearly defined ligand density, **Video 2** shows the waters and ligand density in the binding pocket and compares the deposited to the automated processing maps. Density for water molecules in the ligand binding pocket are clearly resolved in both sets of maps.

**Improvement in Model Building Metrics**

To further confirm the aforementioned improvements in map quality, we remodeled the deposited atomic model against the automated density map for each of the datasets. With reduced anisotropy in the automated maps, we expected there to be less overfitting of the models due to anisotropy induced compression/stretching of the map, resulting in better metrics. After remodeling all datasets, including the apo-CAK dataset, we measured a series of map-model and map based metrics to quantify improvements.

**Map-model metrics show that automated maps yield better models**

Map-model Fourier Shell Correlation is a useful metric to determine how well a model agrees with a map. **Figure 8** shows map-model FSCs for all map-model pairs (excluding apo-CAK, EMPIAR-11800; no model deposited). Map-model FSC curves are considered reliable up to a cutoff of 0.5. Comparing the curves for automated and deposited map-model pairs, there is clear improvement in medium to high resolution ranges for the automated maps compared to the deposited maps, as the FSC curves stay closer to 1.0. This indicates there is more agreement between the automated maps and models compared to the deposited maps/models.

<figure><img src="/files/BymHFFevqYgbBILY84bz" alt=""><figcaption><p><strong>Figure 8. Map-model FSCs for automated and deposited data showing the improvement in information content from DC to about 2.5Å.</strong> Deposited map-model FSCs (orange) and automated map-model FSCs (purple) computed from models refined into the automated maps. Map-model FSCs were computed using Phenix; gray, horizontal, dashed line at FSC = 0.5 for comparison.</p></figcaption></figure>

We also quantified four other map-model parameters to see if other commonly tracked metrics improved. Using the validation tools in Phenix, we measured the CCmask, EMRinger, MolProbity, and Clash scores for all deposited and automated map-model pairs and plotted in **Figure 9** ([Liebschner et al., 2019](https://doi.org/10.1107/S2059798319011471)).

<figure><img src="/files/K0UbI6TGyXT2zgir5rdC" alt=""><figcaption><p><strong>Figure 9. Reduction in map anisotropy improves all measured map-model metrics.</strong> Phenix was used to compute CCmask (left), EMRinger (second column), MolProbity (third column), and Clash scores (right). Deposited map-model metrics are shown with gray circles while automated map-model metrics are shown with filled circles in green (improved), blue (same), and yellow (reduced). Note X-axis units for each plot as the axes are always plotted with improving values to the right.</p></figcaption></figure>

CCmask scores measure the fit of atomic centers to the map, providing a quantitative understanding of model accuracy ([Afonine et al., 2018](https://doi.org/10.1107/S2059798318009324)). Using Phenix, we computed CCmask scores for the deposited and automated map-model pairs across all datasets. In every instance, our automated maps demonstrated improved model fits, with a maximum improvement of 0.16 observed for EMPIAR-11823d.

The EMRinger score evaluates the favorability of side-chain conformations relative to expected rotameric positions. It is quantified as the maximum normalized Z-score (adjusted for model length) across density thresholds, reflecting the agreement between the structural model and the map ([Barad et al., 2015](https://doi.org/10.1038/nmeth.3541)). Higher scores indicate superior model fit. We observed significant EMRinger score improvements ranging from 0.17 to 1.74 (averaging 0.77; n=8) for the automated map-model pairs compared to deposited results.

The MolProbity Score is computed using a “weighted function of clashes, Ramachandran favored, and rotamer outliers, scaled and normalized so that its value approximates the resolution at which that score would be average” ([Williams et al., 2018](https://doi.org/10.1002/pro.3330)). A higher score indicates more clashes, more Ramachandran outliers, and more rotamer outliers. All automated map model pairs have a lower MolProbity score with up to a 36% reduction for EMPIAR-11823c.

The Clash Score, also computed in Phenix, is derived from the MolProbity all-atom steric clash analysis and reported as the number of overlaps ≥0.4Å per thousand atoms ([Afonine et al., 2018](https://doi.org/10.1107/S2059798318009324); [Williams et al., 2018](https://doi.org/10.1002/pro.3330)). Similar to MolProbity, lower values indicate fewer steric conflicts. For all automated datasets, we achieved a reduction in clashes by an average of 3 per thousand atoms; for models containing approximately 10,000 atoms, this corresponds to a total reduction of 30 clashes per model.

#### Per-residue cross-correlation improvements for the three most improved datasets

To further illustrate the map quality improvement on three of the most improved datasets, we evaluated the CCmask values on a more granular level by utilizing both the per-residue values and a 3-residue rolling window across all chains in the model (**Figure 10**): Chain H (Cyclin H), Chain I (MAT1), and Chain J (CDK7). We performed this analysis for EMPIAR entries 11823b, 11823d, and 11823ef. These values were calculated for every residue and water molecule modeled into the respective maps. Across these three datasets, it is evident that the CCmask values are consistently higher for the automated workflow results when compared to the manual processing outcomes. This represents a general increase across the entire map-model pair rather than better resolution in only a single chain or feature.

<figure><img src="/files/MreFxRIU7M6bi0LVTmoi" alt=""><figcaption><p><strong>Figure 10. Cross correlation mask (CCmask) scores for three most improved datasets showing that the improvement is model wide and not due to a specific, local improvement.</strong> Deposited (orange) and automated (purple) CCmask scores on a 3-residue rolling average (dark lines) and per-residue data in translucent lines for each polypeptide (CDK7, MAT1, and Cyclin H) as well as all waters.</p></figcaption></figure>

#### Automated curation isolates smaller, higher quality particle stacks

In most datasets, we observed that final particle counts from automated processing were lower than the deposited manual processing (**Table 2**). This is likely attributable to the rigorous automated curation steps required to eliminate map anisotropy. Our curation strategy prioritizes a strict acceptance of high-quality particles rather than maximizing the total particle count, ultimately yielding better results through more selective filtering.

<table><thead><tr><th width="100.88671875">Dataset</th><th width="219.96484375">Deposited Final # Particles</th><th width="226.69140625">Automated Final # Particles</th></tr></thead><tbody><tr><td>11800</td><td>486,457</td><td>93,732</td></tr><tr><td>11793</td><td>555,117</td><td>468,128</td></tr><tr><td>11799</td><td>126,177</td><td>132,961</td></tr><tr><td>11807</td><td>176,991</td><td>167,148</td></tr><tr><td>11823a</td><td>294,085</td><td>88,447</td></tr><tr><td>11823b</td><td>334,554</td><td>97,538</td></tr><tr><td>11823c</td><td>294,740</td><td>103,098</td></tr><tr><td>11823d</td><td>565,148</td><td>82,165</td></tr><tr><td>11823ef</td><td>269,962</td><td>187,977</td></tr></tbody></table>

<sup>**Table 2 -**</sup> <sup></sup><sup>Particle counts used for the final reconstructions of deposited and automated maps. Particle counts for the automated workflow are smaller due to extensive curation and pruning of over represented views.</sup>

#### End-to-end processing times less than 24 hours on average

All processing for these datasets was completed on a single GPU node utilizing 32 CPU cores, 512 GB of RAM, and eight NVIDIA A100 (40GB) GPUs. We utilized parallelization across resources where possible; Reference Based Motion Correction (RBMC) which was restricted to 6 GPUs. The reported runtimes in **Table 3** represent end-to-end processing for the entire workflow, inclusive of RBMC. For EMPIAR-11823e and 11823f, particle stacks were initially processed separately (requiring 16.4 and 16.6 hours, respectively) before being merged. We anticipate that total processing time for these combined datasets would align with 11793 or 11807 given the similar movie and particle counts. Average processing times across all datasets were approximately 21 hours.

Ultimately, for datasets collected in an unattended manner, by leveraging CryoSPARC Live for real-time preprocessing during collection and employing parallel processing strategies with automated workflows in CryoSPARC, it could be feasible to collect and process a fully loaded 12-grid autoloader within a handful of days.

<table><thead><tr><th width="111.0625">Dataset</th><th width="132.45703125"># of movies</th><th width="216.67578125">Total processing time (hr)</th></tr></thead><tbody><tr><td>11800</td><td>5,805</td><td>17.9</td></tr><tr><td>11793</td><td>9,118</td><td>23.9</td></tr><tr><td>11799</td><td>5,780</td><td>17.5</td></tr><tr><td>11807</td><td>10,017</td><td>25.2</td></tr><tr><td>11823a</td><td>5,729</td><td>18.7</td></tr><tr><td>11823b</td><td>5,583</td><td>18.3</td></tr><tr><td>11823c</td><td>5,210</td><td>17.0</td></tr><tr><td>11823d</td><td>5,203</td><td>18.0</td></tr><tr><td>11823ef</td><td>10,598</td><td>33.0</td></tr></tbody></table>

<sup>**Table 3**</sup> <sup></sup><sup>- Number of movies and total processing time (hours) for all datasets interrogated in this study. These timings include all steps in the workflow.</sup>

{% hint style="success" icon="arrow-up-left" %}
Maps and fitted models from the results of the automated data processing can be found here: [CAK\_automation\_maps\_models.zip](https://structura-assets.s3.us-east-1.amazonaws.com/automated-workflows/CAK_automation_maps_models.zip)
{% endhint %}

## How we iteratively developed an improved automation workflow to handle preferred orientation

Processing the apo-CAK dataset was undertaken with the explicit goals of understanding any data processing intricacies or pitfalls related to this target and generating a reference volume for further processing of drug bound CAK datasets. First, [Workflow v1](https://guide.cryosparc.com/processing-data/automated-workflows/preprint-end-to-end-fully-automated-data-processing-of-gpcrs) was applied, but the jobs were not queued (**Figure 18** - column 1). Each job in the workflow was manually run to process the apo-CAK dataset, allowing for inspection of results and to ensure parameters produce the desired outcomes before proceeding to the next job. Below, we explore these results along with observations made throughout processing that lead to an optimized workflow (**Figure 18** - column 4) for automated processing of the other CAK datasets.

### Initial workflow

#### Preprocessing

Micrograph preprocessing steps including [Import Movies](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/import/job-import-movies#imported-movies), Patch Motion Correction, Patch CTF Estimation, [Manually Curate Exposures](https://guide.cryosparc.com/application-guide/interactive-jobs#:~:text=or%20similar%20particle.-,Interactive%20Job%3A%20Manually%20Curate%20Exposures,-Manually%20curate%20exposures), Micrograph Junk Detector, and Micrograph Denoiser. Similar to the GPCR workflow, micrographs with a CTF fit resolution worse than 5Å and total full-frame motion greater than 200 pixels were excluded. This resulted in the rejection of 116 micrographs; leaving behind \~5,700 for the next processing steps. Micrographs put through this preprocessing pipeline result go from low contrast and difficult to interpret (**Figure 11A**) to easy to visualize particles and labeled junk (**Figure 11B**).

<div><figure><img src="/files/RxvSEFlohr4utQI6UxvB" alt="" width="375"><figcaption><p><strong>Figure 11A. Motion corrected micrograph of EMPIAR-11800 showing poor SNR for particles and a large swath of ethane contamination.</strong></p></figcaption></figure> <figure><img src="/files/fqcs81nwYajM2SXGAFPb" alt="" width="375"><figcaption><p><strong>Figure 11B. The same micrograph as the left after junk is labeled and the micrograph is denoised and ready for particle picking.</strong></p></figcaption></figure></div>

#### Particle Picking

Having preprocessed micrographs, the [Blob Picker](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/particle-picking/job-blob-picker#circular-blob) was used with a minimum particle diameter of 60Å, a maximum diameter of 90Å, a minimum separation of 0.7 particle diameters to pick particles from 400 denoised micrographs. Following picking, particles were extracted with a box size of 300 pixels Fourier cropped to 128 pixels. Following blob picking, particle picks were filtered by first the Micrograph Junk Detector to remove particle picks on junk and then by the auto mode of Inspect Particle Picks. The power score for auto Inspect Particle Picks had to be reduced to 20 due to the smaller size of the particle and lack of a micelle compared to GPCRs. Particles were then subject to [2D Classification](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/particle-curation/job-2d-classification#number-of-2d-classes) to generate templates for template picking.

<figure><img src="/files/eycnjF1x3wYy2T7AN33b" alt=""><figcaption><p><strong>Figure 12. Poor orientation distribution in 2D-classification utilizing blob picking.</strong> 2D classes of particles picked from 400 micrographs.</p></figcaption></figure>

Right away there were two things to note about some classes in **Figure 12**:

1. Classes that were higher resolution and had clear secondary structural features represented a very similar viewing direction of the CAK complex.
2. A few classes at the bottom indicated there were hot pixels present. After some investigation we found that these came from a defective gain file. Outliers in the gain file were set as defective pixels and movies were re-motion corrected before continuing with processing.

Processing was continued at this point where good templates, **Figure 13**, were selected from the above 2D classification to be used for template picking. Although we recognized this would likely lead to a poor reconstruction, this was done to demonstrate what would happen if poor templates were used to pick particles. These templates were used to template pick all micrographs in the dataset. Once all micrographs were picked, particle picks were put through the Micrograph Junk Detector and the automated modes of Inspect Particle Picks to remove picks on junk or on the background resulting in a particle stack that was roughly 2.1M particles. These particles were extracted from the micrographs in a box size of 384 pixels and Fourier cropped to 128 pixels.

<figure><img src="/files/JQvvtT8uE9iOgA4Sw5mN" alt=""><figcaption><p><strong>Figure 13. Orientation bias is introduced when selecting templates from the above 2D-classification job.</strong> Classes selected from 2D Classification of blob picked particles. Notable, only a few views, mostly of the broad side of the CAK complex.</p></figcaption></figure>

#### Particle Curation

Since we were processing the first dataset in the series (apo-CAK), we did not have a reference volume to use in a decoy classification. Therefore, the cleaned particle stack was first subjected to multi-class Ab-Initio Reconstruction to generate a set of volumes to use for heterogeneous refinement. For this Ab-Initio Reconstruction, the initial resolution and maximum resolution are increased to 20Å and 5Å, respectively. These resolution choices were inspired by a recent preprint about High-resolution Ab-Initio Reconstruction (HR-HAIR; [Kim et al., 2025](https://doi.org/10.1101/2025.09.08.674935)) and our development of [Homogeneous Ab-Initio Refinement in CryoSPARC v5](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/3d-reconstruction/job-homogeneous-ab-initio-refinement-beta) . We found that to get a successful volume that would lead to desirable results from Heterogeneous Refinement, the 5Å maximum resolution was key. Even decreasing the maximum resolution to 8Å resulted in unusable volumes. Full parameter selections for Ab-Initio Reconstruction are below, and these were used for all Ab-Initio Reconstruction jobs moving forward unless otherwise noted.

<table><thead><tr><th width="257.171875">Parameter</th><th>Setting</th><th>Explanation</th></tr></thead><tbody><tr><td><code>Number of Ab-Initio classes</code></td><td>3</td><td>3 classes seemed appropriate since the sample was relatively pure and there were few components present</td></tr><tr><td><code>Number of particles to use</code></td><td>250,000</td><td>Using 10% of the initial particle stack usually results in a decent set of volumes</td></tr><tr><td><code>Maximum resolution</code></td><td>5</td><td>Smaller particles typically benefit from higher resolutions - generating better volumes</td></tr><tr><td><code>Initial Resolution</code></td><td>20</td><td></td></tr></tbody></table>

Following multi-class Ab-Initio Reconstruction, all particles and volumes were subjected to Heterogeneous Refinement with default parameters. **Figure 14** shows one class (green) depicting CAK that was manually selected to move forward with in processing, but it is highly anisotropic.

<figure><img src="/files/UDcxqn5jUsYj89zhbMS5" alt=""><figcaption><p><strong>Figure 14. Initial volumes after the first Heterogeneous Refinement, displayed at the same threshold standard deviation.</strong> The green volume appears to be the shape of CAK, but is extremely anisotropic.</p></figcaption></figure>

Following this, duplicates were pruned from the particle stack of the selected green class, and then were re-extracted at a box size of 384 pixels and Fourier cropped to 256 pixels. These particles were then subjected to a second multi-class Ab-Initio Reconstruction using the following parameters:

<table><thead><tr><th width="280.265625">Parameter</th><th width="103.78515625">Setting</th><th>Explanation</th></tr></thead><tbody><tr><td><code>Number of Ab-Initio classes</code></td><td>5</td><td>Since there is significant anisotropy, opted for more classes</td></tr><tr><td><code>Number of particles to use</code></td><td>100,000</td><td>Using ~10% of the initial particle stack which usually results in a decent set of volumes</td></tr><tr><td><code>Maximum resolution</code></td><td>5</td><td>Smaller particles typically benefit from higher resolutions - generating better volumes</td></tr><tr><td><code>Initial Resolution</code></td><td>10</td><td>Inspired by the HR-HAIR preprint, we opted to begin perform pose estimation at medium resolutions</td></tr></tbody></table>

All particles and volumes from the Ab-Initio Reconstruction were used for another Heterogeneous Refinement using the following settings:

<table><thead><tr><th width="215.21875">Parameter</th><th width="94.40625">Setting</th><th>Explanation</th></tr></thead><tbody><tr><td><code>Refinement box size</code></td><td>256</td><td>Using a larger box size allows higher resolution reconstructions</td></tr><tr><td><code>Batch size per class</code></td><td>3000</td><td>Increasing this number can provide better classification results and this value was used in the GPCR workflow</td></tr><tr><td><code>Initial Resolution</code></td><td>12</td><td>Lower resolutions lead to poorer results in initial iterations</td></tr></tbody></table>

The best class (**Figure 15**: green) was selected for further processing. Since there was only one good class, the particles associated with this class were reconstructed with a Homogeneous Reconstruction job and refined using Non-uniform Refinement. After Non-uniform Refinement the map exhibited anisotropy, and the density in many regions was discontinuous (**Figure 19**, column 1). Additionally, the low cFAR score (0.09) matched the quality of the reconstruction.

<figure><img src="/files/DI3apTdYP0Nxlobqsrqo" alt=""><figcaption><p><strong>Figure 15. Volumes from second round of particle curation (heterogeneous refinement), with the green being the good volume that moved through the processing pipeline.</strong> Green volume was selected for further processing. Other volumes present, notable class 0 and class 3 (left to right) look like CAK, but are highly anisotropic.</p></figcaption></figure>

### Addressing preferred orientation

At this point, all jobs related to particle curation in the [Workflow v1](https://guide.cryosparc.com/processing-data/automated-workflows/preprint-end-to-end-fully-automated-data-processing-of-gpcrs) had been successfully run, but a low cFAR score indicated there was remaining map anisotropy that needed to be addressed. First, the Rebalance Orientations job type was used to remove particles from viewing direction bins that were over-represented. The Non-uniform Refinement leading into the Rebalance Orientations job showed a slight shoulder of low scale particles emerging and thus overrepresented particles with the lowest per-particle scale (PPS) values were removed from their viewing direction bin. Although it appears that many particles were removed with scales greater than 1.0, it should be noted that the particles being removed are those that are overrepresented and thus would be considered “good particles”, but in comparison to their binned counterparts, have a lower per-particle scale. Since the resulting rebalanced particle stack has a clear, bi-modal distribution (**Figure 16C,D**), Subset Particles by Statistic was used to to remove low scale particles, and retaining particles in Cluster 1 (light blue) for further processing.

<figure><img src="/files/lfdS050aDGBoSyD8o4Wu" alt=""><figcaption><p><strong>Figure 16. Rebalanced orientations and per-particle scale values for green map of Figure 14 showing over-represented views that can be pruned along with the low-scale particles.</strong> (A,B) Rebalancing the orientations reduced the number of over represented views present in the dataset. (C) After rebalancing orientations, a bimodal distribution of per-particle scale values are observed (blue) once the over represented views (red) are removed. (D) Bi-modal distribution of CAK per-particle scale values, clustered into low scale (dark blue) and high scale (light blue) by the job Subset Particles by Statistic.</p></figcaption></figure>

The resulting particle stack was then subjected to a final Homogeneous Reconstruction and Non-Uniform Refinement. For the Non-Uniform Refinement, global (3rd order only) and local CTF corrections were enabled. The resulting volume had less visible anisotropy and the cFAR score improved from 0.09 to 0.31 (**Figure 19,** column 2). Although this map looked better than the first, further approaches were undertaken with changes to particle picking and curation steps with hopes to reduce anisotropy and improve the connectivity of map density.

### Intermediate workflow changes

Still processing in a semi-manual manner (selecting volumes manually to move through the workflow), we aimed to improve the resulting map anisotropy. We hypothesized that changing the picking strategy might pick more rare views, especially since the templates we used in the first workflow were all of one view of the CAK complex. Therefore, using the Blob Picker with the following parameters, picking was carried out on all micrographs using circular blobs.

<table><thead><tr><th width="312.984375">Parameter</th><th width="97.953125">Setting</th><th width="422.78125">Explanation</th></tr></thead><tbody><tr><td><code>Minimum particle diameter (Å)</code></td><td>40</td><td>Based on the previous reconstruction, this could be the smallest view</td></tr><tr><td><code>Maximum particle diameter (Å)</code></td><td>90</td><td>Based on the previous reconstruction, this could be the largest view</td></tr><tr><td><code>Pick on denoised micrographs</code></td><td>True</td><td>Picking on denoised micrographs yields better picks and allows filtering via auto-Inspect Particle Picks</td></tr><tr><td><code>Min. separation dist (diameters)</code></td><td>0.6</td><td>Likely too low, but the results of the picking job looked good upon manual inspection</td></tr></tbody></table>

After particle picking, picks were cleaned as before using the Micrograph Junk Detector and Auto-Inspect Particle Picks. Next, a similar set of particle curation steps were utilized with a few changes. The first change was using a funnel-like approach to filtering poor particles where the first multi-class Ab-Initio Reconstruction and Heterogeneous Refinement used five classes and the second round used three classes. Between the two rounds of 3D particle curation, another Rebalance Orientations job was included to prune out overrepresented views early on. This helped to reduce anisotropy before the second round of particle classification.

Looking to the effects of these above changes, **Figure 19** (column 3) shows that the effects of these changes were minimal, but in the correct direction. Resolution and cFAR scores improved but the quality of density was roughly the same.

### Final Workflow Changes: Workflow v2

As a final approach to reducing anisotropy and improving the density of the final reconstruction, we opted for a different particle picking strategy and initial curation strategy. Both of these improvements utilized the apo-CAK volume that was output at the end of intermediate workflow. This volume was lowpass filtered to 15Å to reduce any high resolution bias for the two major changes to the workflow.

Using the apo-CAK volume output at the end of the intermediate workflow, the Create Templates job with default parameters was used to create 50 templates, uniformly spaced on the viewing sphere, to use for particle picking (**Figure 17**).

<figure><img src="/files/FgdvnwveArHzkWw3YhK0" alt=""><figcaption><p><strong>Figure 17. Templates created from LPF apo-CAK volume output from the second workflow.</strong></p></figcaption></figure>

These templates were used with the Template Picker on denoised micrographs and the below parameters to fully pick all micrographs. After particle picking, picks were cleaned up using the Micrograph Junk Detector and Auto-Inspect Particle Picks.

<table><thead><tr><th width="305.51953125">Parameter</th><th width="100.22265625">Setting</th><th width="465.60546875">Explanation</th></tr></thead><tbody><tr><td><code>Particle diameter (Å)</code></td><td>90</td><td>Based on the previous reconstruction, this could be the largest view</td></tr><tr><td><code>Pick on denoised micrographs</code></td><td>True</td><td>Picking on denoised micrographs yields better picks and allows filtering via auto-Inspect Particle Picks</td></tr><tr><td><code>Use CTFs to filter templates</code></td><td>False</td><td>Automatically disabled when picking on denoised micrographs</td></tr><tr><td><code>Min. separation dist (diameters)</code></td><td>0.6</td><td>Provides a minimum separation distance of 54Å</td></tr></tbody></table>

To help ensure that rare views were correctly classified into a good class during particle curation, we changed the first round of particle curation from Ab-Initio Reconstruction and Heterogeneous Refinement to “decoy classification”. Decoy classification is the use of known volumes (one representing the target class plus a number of junk volumes to sequester poor particles) as inputs into a Heterogeneous Refinement job ([see here for details](https://www.biorxiv.org/content/10.1101/2025.10.17.682689v1.full#:~:text=7.3.1%20Decoy%20Classification)). Decoy classification is useful for (a) providing a reference accurate enough to capture all relevant views of the particle and (b) easily enabling automated selection of good particles (as they presumably end up in the good class). The low resolution reference of apo-CAK lowpass filtered to 15Å was used as the good volume and the same junk volumes from the GPCR datasets (see [Table 1 of the preprint](https://doi.org/10.1101/2025.10.17.682689)) were used for decoy classification with the apo-CAK particle stack.

Following decoy classification, particles in the good class were globally refined to obtain more accurate pose information and to compute per-particle scale values before pruning over-represented views with Rebalance Orientations. As is the same in intermediate workflow, particles were then subjected to a second round of curation using Ab-Initio Reconstruction and Heterogeneous Refinement with the same parameter settings as previous workflows. Use of Reference Based Auto Select 3D after the second round of curation allowed for automatic selection of the good class(es) based on resolution. Particles from the high-resolution, good classes were then reconstructed, globally refined, and particles from over-represented views and particles with low per-particle scale values were removed before performing Reference Based Motion Correction and a final Non-uniform Refinement.

After implementing these changes, the workflow was run against the apo-CAK dataset. The resulting volume had more connectivity, achieved higher resolution, and the cFAR score was further improved from the intermediate workflow, 0.35 to 0.40 (**Figure 18** - Final Workflow, **Figure 18** - column 4).

### Workflow evolution

**Figure 18** details all incremental changes made to Workflow v1, which ultimately yielded Workflow v2. In totality, the particle picking stage went through the most changes and the particle curation stages feature the inclusion of two Rebalance Orientations job and one Subset Particles by Statistic.

<figure><img src="/files/bOCLDrSAp1XUZqrduvYD" alt=""><figcaption><p><strong>Figure 18. Workflow evolution across trials</strong> - (Column 1) Workflow v1 applied to CAK. (Column 2) Small addition to Workflow v1 to reduce anisotropy encountered in the CAK datasets. (Column 3) Second workflow trying to improve map anisotropy. (Column 4) Final workflow settled upon for automated data processing of CAK (Workflow v2). Jobs with black outlines constitute changes from the previous column.</p></figcaption></figure>

The changes to the workflow can be directly visualized in **Figure 19**, wherein the map anisotropy is reduced at each version of the workflow as judged by the cFAR scores. In addition, the map density displays more connectivity in the final volume of the workflow in comparison to the first.

<figure><img src="/files/VrfA8fppuePz9OkbTSi7" alt=""><figcaption><p><strong>Figure 19. Clear improvements in the map quality and cFAR score between workflows.</strong> GSFSC curves and cFSCs for all four volumes were computed with the same mask for fair comparison.</p></figcaption></figure>

## Conclusion

This case study demonstrates that automation in CryoSPARC is robust and generalizable, and can produce results that are equal or better than manual processing, even on a challenging series of datasets of a small drug-bound particle with strong preferred orientation. The updated Workflow v2 used here can be downloaded and easily used to automate processing of your own targets.

Automation of data processing in CryoSPARC enables significant reductions in time-to-structure, particularly for SBDD scenarios. Processing times on the order of 20 hours can be obtained on modest, local compute infrastructure. Furthermore, for a scientist to adapt our published workflows to their own target, the time investment is minimal, typically requiring only one or two days of exploratory processing with an exemplar dataset. After this initial work, all subsequent datasets of the same target can be processed fully automatically.

## How to set up automation in CryoSPARC for a new target class

All of the methods and tools required to set up an automated workflow for a new target class, such as CAK, are already available in CryoSPARC v4.7.1+. A workflow file, reference volume, final maps, and final models from the automation study are available here:&#x20;

{% hint style="success" icon="arrow-up-left" %}
[automated\_workflow\_materials\_v2.zip](https://structura-assets.s3.us-east-1.amazonaws.com/automated-workflows/Automated_workflow_materials_v2.zip)
{% endhint %}

To apply the Workflow to a new target, you will need a reference volume of your target from a previously processed dataset. Granted these items are available, you can:

1. Download the workflow JSON files and junk volumes: [automated\_workflow\_materials\_v2.zip](https://structura-assets.s3.us-east-1.amazonaws.com/automated-workflows/Automated_workflow_materials_v2.zip)
2. [Upload these volumes](https://guide.cryosparc.com/processing-data/automated-workflows/practical-tips-uploading-files-and-using-workflows#uploading-volumes-and-masks-for-use-in-workflow) to your compute setup.
3. [Import workflow JSON file](https://guide.cryosparc.com/processing-data/automated-workflows/practical-tips-uploading-files-and-using-workflows#importing-the-workflow-json-file) into CryoSPARC using the workflow panel in the sidebar.
4. [Select the workflow from the list](https://guide.cryosparc.com/processing-data/automated-workflows/practical-tips-uploading-files-and-using-workflows#applying-the-workflow) and set all flagged parameters. Parameters that might need to be adjusted (on a per dataset basis) are below:

   Import movies

   * Movies data path
   * Gain reference path
   * Raw pixel size (Å)
   * Accelerating voltage (kV)
   * Spherical aberration (mm)
   * Total exposure dose (e/Å^2)
   * Flip gain ref and defect file in Y
   * Exposure group metadata (if present)

   Import 3D Volumes

   * Path to reference and junk volumes

   Patch Motion Correction

   * Output F-crop factor

   Template Picker

   * Particle diameter (Å) - set to the largest particle dimension

   Extract from micrographs (2x)

   * Extraction box size - set to a value that is 2-3x the size of your target
5. Click on the green “Apply” button in the bottom of the workflow GUI.
6. Queue and run each job independently, or in small clusters. Review all results to ensure that desired outputs are generated and there are no processing issues. A few key results to check include:
   1. Template Picker - ensure minimum separation distance setting resulted in adequate picking (not to sparse, not double picking particles)
   2. Inspect Particle Picks - ensure that the power value selected is correct for the peak of particles ([see GPCR automation preprint for more info](https://www.biorxiv.org/content/10.1101/2025.10.17.682689v1.full#:~:text=7.2.3%20Inspect%20Particle%20Picks%20Auto%20Cluster%20Mode))
   3. Rebalance Orientations - on a target that exhibits little or no preferred orientation, Rebalance Percentile can be increased to a value of 80-90 so as to preserve more of the picked particles.
   4. Extract from Micrographs (2nd instance) - modulate the Fourier crop box size parameter to ensure that the resolution of reconstruction is not inhibited by Nyquist ([more on this here](https://guide.cryosparc.com/processing-data/tutorials-and-case-studies/case-study-dktx-bound-trpv1-empiar-10059#extract-from-micrographs:~:text=the%20particle%20stack.-,Extract%20from%20Micrographs,-In%20the%20early)).
7. If issues arise, add exploratory processing steps to troubleshoot these issues.
8. Once all processing is completed to a satisfactory result, select all jobs, rebuild the workflow, and apply it to other datasets of the same target.

{% hint style="info" %}
If you collected data using beam image-shift, include an Exposure Group Utilities job between Manually Curate Exposures and the Micrograph Denoiser.
{% endhint %}

## References

[Cushing, V. I. ](https://doi.org/10.1038/s41467-024-46375-9)[*et al.*](https://doi.org/10.1038/s41467-024-46375-9)[ ](https://doi.org/10.1038/s41467-024-46375-9)[**High-resolution cryo-EM of the human CDK-activating kinase for structure-based drug design.**](https://doi.org/10.1038/s41467-024-46375-9)[ ](https://doi.org/10.1038/s41467-024-46375-9)[*Nat Commun*](https://doi.org/10.1038/s41467-024-46375-9)[ ](https://doi.org/10.1038/s41467-024-46375-9)[**15**](https://doi.org/10.1038/s41467-024-46375-9)[, 2265 (2024).](https://doi.org/10.1038/s41467-024-46375-9)

[Liebschner, D. ](https://doi.org/10.1107/S2059798319011471)[*et al.*](https://doi.org/10.1107/S2059798319011471)[ ](https://doi.org/10.1107/S2059798319011471)[**Macromolecular structure determination using X-rays, neutrons and electrons: recent developments in** ](https://doi.org/10.1107/S2059798319011471)[***Phenix***](https://doi.org/10.1107/S2059798319011471)[. ](https://doi.org/10.1107/S2059798319011471)[*Acta Crystallogr D Struct Biol*](https://doi.org/10.1107/S2059798319011471)[ ](https://doi.org/10.1107/S2059798319011471)[**75**](https://doi.org/10.1107/S2059798319011471)[, 861–877 (2019)](https://doi.org/10.1107/S2059798319011471)

[Afonine, P. V. ](https://doi.org/10.1107/S2059798318009324)[*et al.*](https://doi.org/10.1107/S2059798318009324)[ ](https://doi.org/10.1107/S2059798318009324)[**New tools for the analysis and validation of cryo-EM maps and atomic models.**](https://doi.org/10.1107/S2059798318009324)[ ](https://doi.org/10.1107/S2059798318009324)[*Acta Crystallogr D Struct Biol*](https://doi.org/10.1107/S2059798318009324)[ ](https://doi.org/10.1107/S2059798318009324)[**74**](https://doi.org/10.1107/S2059798318009324)[, 814–840 (2018).](https://doi.org/10.1107/S2059798318009324)

[Barad, B. A. ](https://doi.org/10.1038/nmeth.3541)[*et al.*](https://doi.org/10.1038/nmeth.3541)[ ](https://doi.org/10.1038/nmeth.3541)[**EMRinger: side chain–directed model and map validation for 3D cryo-electron microscopy.**](https://doi.org/10.1038/nmeth.3541)[ ](https://doi.org/10.1038/nmeth.3541)[*Nat Methods*](https://doi.org/10.1038/nmeth.3541)[ ](https://doi.org/10.1038/nmeth.3541)[**12**](https://doi.org/10.1038/nmeth.3541)[, 943–946 (2015).](https://doi.org/10.1038/nmeth.3541)

[Williams, C. J. ](https://doi.org/10.1002/pro.3330)[*et al.*](https://doi.org/10.1002/pro.3330)[ ](https://doi.org/10.1002/pro.3330)[**MolProbity: More and better reference data for improved all‐atom structure validation.**](https://doi.org/10.1002/pro.3330)[ ](https://doi.org/10.1002/pro.3330)[*Protein Science*](https://doi.org/10.1002/pro.3330)[ ](https://doi.org/10.1002/pro.3330)[**27**](https://doi.org/10.1002/pro.3330)[, 293–315 (2018).](https://doi.org/10.1002/pro.3330)

[Kim, K., Li, H. & Clarke, O. B. ](https://doi.org/10.1101/2025.09.08.674935)[**High-resolution ab initio reconstruction enables cryo-EM structure determination of small particles.**](https://doi.org/10.1101/2025.09.08.674935)[ ](https://doi.org/10.1101/2025.09.08.674935)[*bioRxiv* ](https://doi.org/10.1101/2025.09.08.674935)[(2025)](https://doi.org/10.1101/2025.09.08.674935).
