# About CryoSPARC™

General information about the software platform.

## What is CryoSPARC™?

CryoSPARC is a state of the art scientific software platform for cryo-electron microscopy (cryo-EM) used in research and drug discovery pipelines. CryoSPARC is used to reconstruct and visualize cryo-EM structures of biological targets including membrane proteins, viruses and complexes, from raw movies to high resolution maps. Learn more: <https://cryosparc.com/>

## What is CryoSPARC Live™?

CryoSPARC Live is an extension of CryoSPARC that delivers real-time cryo-EM data processing, quality assessment and feedback as data is collected. Learn more: <https://cryosparc.com/live>

## Licensing

CryoSPARC and CryoSPARC Live can be licensed for non-profit academic use and commercial use. Please see:

{% content-ref url="/pages/-M8MGhEyJ3OFedB7vEjq" %}
[Licensing](/licensing)
{% endcontent-ref %}

## Citation

If you use CryoSPARC in your work, please cite as follows.

### CryoSPARC algorithms

Please cite the following papers as appropriate:

* **General CryoSPARC/CryoSPARC Live use, including preprocessing, 2D classification, Ab-initio reconstruction, Refinement:**\
  [Punjani, A., Rubinstein, J.L., Fleet, D.J. & Brubaker, M.A. cryoSPARC: algorithms for rapid unsupervised cryo-EM structure determination. Nature Methods 14, 290-296 (2017).](https://www.nature.com/articles/nmeth.4169)
* **Local (per-particle) motion correction:**\
  [Rubinstein, J.L. & Brubaker, M.A. Alignment of cryo-EM movies of individual particles by optimization of image translations. Journal of Structural Biology 192 (2), 188-195 (2015).](https://www.sciencedirect.com/science/article/pii/S1047847715300459)
* **Non-uniform refinement:**\
  [Punjani, A., Zhang, H. & Fleet, D.J. Non-uniform refinement: adaptive regularization improves single-particle cryo-EM reconstruction. Nat Methods **17,** 1214–1221 (2020).](https://www.nature.com/articles/s41592-020-00990-8)
* **3D Variability Analysis:**\
  [Punjani, A. & Fleet, D.J. 3D variability analysis: Resolving continuous flexibility and discrete heterogeneity from single particle cryo-EM. Journal of Structural Biology, Volume 213, Issue 2, 2021.\
  https://doi.org/10.1016/j.jsb.2021.107702](https://doi.org/10.1016/j.jsb.2021.107702)
* **3D Flexible Refinement:**\
  [Punjani, A. & Fleet, D.J. 3DFlex: determining structure and motion of flexible proteins from cryo-EM. Nature Methods (2023). https://doi.org/10.1038/s41592-023-01853-8](https://doi.org/10.1038/s41592-023-01853-8)

### CryoSPARC implementations

* **ResLog analysis:**\
  [Stagg, S.M., Noble, A.J., Spilman, M. & Chapman, M.S. ResLog plots as an empirical metric of the quality of cryo-EM reconstructions. Journal of Structural Biology 185 (3), 418-426 (2014).](https://www.sciencedirect.com/science/article/pii/S1047847713003377?via%3Dihub)
* **CTF refinement and aberration correction:**\
  [Zivanov, J., Nakane, T. & Scheres, S. H. W. Estimation of high-order aberrations and anisotropic magnification from cryo-EM data sets in *RELION*-3.1. IUCrJ 7, 253-267 (2020).](https://dx.doi.org/10.1107%2FS2052252520000081)
* **Reference-based motion correction/Bayesian polishing:**\
  [Zivanov, J., Nakane, T. & Scheres, S. H. W. A Bayesian approach to beam-induced motion correction in cryo-EM single-particle analysis. IUCrJ 6, 5-17 (2019).](https://doi.org/10.1107/S205225251801463X)

### Wrappers to third-party tools

Users should obtain their own software licenses (as applicable) for the below programs, for which wrappers are available in CryoSPARC.

* **MotionCor2:** Shawn Q. Zheng, Eugene Palovcak, Jean-Paul Armache, Yifan Cheng and David A. Agard (2016) Anisotropic Correction of Beam-induced Motion for Improved Single-particle Electron Cryo-microscopy, Nature Methods, submitted. BioArxiv: <http://biorxiv.org/content/early/2016/07/04/061960>
* **CTFFIND:** [Rohou, A. & Grigorieff, N. CTFFIND4: Fast and accurate defocus estimation from electron micrographs. Journal of Structural Biology 192 (2), 216-221 (2015).](https://www.sciencedirect.com/science/article/pii/S1047847715300460?via%3Dihub)
* **Gctf:** Gctf: Jack (Kai) Zhang. Zhang, K. (2016). Gctf : Real-time CTF determination and correction. Journal of Structural Biology, 193(1), 1-12. <https://doi.org/10.1016/j.jsb.2015.11.003>
* **Topaz:** Bepler, T., Morin, A., Rapp, M. et al. Positive-unlabeled convolutional neural networks for particle picking in cryo-electron micrographs. Nat Methods 16, 1153–1160 (2019) doi:10.1038/s41592-019-0575-8 and Bepler, T., Noble, A.J., Berger, B. Topaz-Denoise: general deep denoising models for cryoEM. bioRxiv 838920 (2019) doi: <https://doi.org/10.1101/838920>
* **3DFSC:** [Tan, Y.Z., Baldwin, P.R., Davis, J.H., Williamson, J.R., Potter, C.S., Carragher, B. & Lyumkis, D. Addressing preferred specimen orientation in single-particle cryo-EM through tilting. Nature Methods 14, 793-796 (2017).](http://dx.doi.org/10.1038/nmeth.4347)
* **DeepEMhancer:** R. Sanchez-Garcia, J. Gomez-Blanco, A. Cuervo et al., “DeepEMhancer: a deep learning solution for cryo-EM volume post-processing”, Communications Biology, vol. 4, no. 874, 2021. Available: 10.1038/s42003-021-02399-1.

### Dependencies

#### **cuDNN**

libcudnn.so.8 is distributed with CryoSPARC as of v3.2, pursuant to the terms of NVIDIA's Software License Agreement (SLA) for cuDNN: <https://docs.nvidia.com/deeplearning/cudnn/sla/index.html>

#### **scikit-cuda**

A modified version of scikit-cuda is included with cryosparc\_compute as of v3.2, pursuant to the scikit-cuda license terms: <https://scikit-cuda.readthedocs.io/en/latest/>

> Redistributions of source code must retain the above copyright notice, this list of conditions and the following disclaimer. Redistributions in binary form must reproduce the above copyright notice, this list of conditions and the following disclaimer in the documentation and/or other materials provided with the distribution. Neither the name of Lev E. Givon nor the names of any contributors may be used to endorse or promote products derived from this software without specific prior written permission.
>
> THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS “AS IS” AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR CONTRIBUTORS BE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.

## CryoSPARC in Scientific Studies

Hundreds of structural studies have used CryoSPARC for cryo-EM data processing:

{% hint style="success" %}
[Google Scholar: cryoSPARC](https://scholar.google.ca/scholar?cites=6690181732944497496\&as_sdt=2005\&sciodt=0,5\&hl=en)
{% endhint %}

## Development

CryoSPARC was originally a research project with origins at the University of Toronto in 2014. As of 2016, all research and development for CryoSPARC is done by [Structura Biotechnology Inc.](https://structura.bio/), a scientific software company based in Toronto, Canada.

By combining our expertise in image processing, algorithm development and professional software engineering, we aim to keep CryoSPARC at the forefront of software for cryo-EM. To that end, we are constantly working on new algorithms and software features which we release on an ongoing basis. CryoSPARC's GPU-accelerated code is written entirely from scratch in-house, with exception of certain wrappers to third party tools that are clearly indicated in the documentation. Many of the algorithms in CryoSPARC are novel developments for cryo-EM image processing and links to publications can be found throughout this documentation.

### Major Version History

* CryoSPARC v5.0 was released on January 27, 2026.
* CryoSPARC v4.0 was released on October 3, 2022 and has been followed by several subsequent releases up to v4.7.1.
* CryoSPARC v3.0 was released on December 9, 2020 and has been followed by subsequent version v3.1, v3.2 and v3.3.
* CryoSPARC v2.0 (released August 17, 2018) was followed by a number of new releases up to v2.15.0 (released May 13, 2020).
* CryoSPARC v0.2.1 was the first public version of CryoSPARC (released February 7, 2017) and was followed by a number of new releases up to v0.6.5 (released January 12, 2018).

For release notes, see: <https://cryosparc.com/updates>

© 2026 Structura Biotechnology Inc. All rights reserved.\
CryoSPARC™ and CryoSPARC Live™ are trademarks of Structura Biotechnology Inc.


# Licensing

Non-profit and commercial licensing options.

## Non-profit use

CryoSPARC™ and CryoSPARC Live™ are available free of charge for non-profit academic research. Non-Profit Academic Research is defined as follows:

{% hint style="info" %}
**Non-Profit Academic Research** means practicing, making, using, improving upon, importing and exporting (but not selling, leasing or otherwise monetizing) academic or scholarly research, for individual (personal) or academic institutional research purposes, in good faith, and expressly excludes, without limitation, purposes that are intended to (or result in, whether by intent or otherwise): (i) create a commercial advantage for any Person; (ii) generate monetary compensation for products or services; (iii) generate commercialization rights for any Person; (iv) be used in an ongoing business concern; or (v) result in an ongoing business concern obtaining any intellectual property rights in any research or results linked to the Non-profit Academic Research.
{% endhint %}

To request a license for non-profit academic research, please fill out the form on our website:

{% embed url="<https://cryosparc.com/download>" %}

The full text of the non-commercial license agreement is available here for reference:

{% content-ref url="/pages/-M8MHBdk2uDzL7ALZQI2" %}
[Non-commercial license agreement](/licensing/non-commercial-license-agreement)
{% endcontent-ref %}

## For-profit, commercial or industry use

Please contact <sales@structura.bio> for inquiries relating to for-profit or industry licensing, including academic-industry collaborations on proprietary projects and fee-for-service data processing.

## Questions about licensing?

If you aren't sure which license applies to your use case, or have any questions around licensing, please contact us at <sales@structura.bio> with your questions.


# Non-commercial license agreement

Full text of CryoSPARC non-commercial license agreement. Please send any queries to: info\@structura.bio.

## CryoSPARC Non-Commercial Software License Agreement

*Last Updated May 27, 2020*

This Non-Commercial Software License Agreement (the "Agreement") is made between you (the "Licensee") and Structura Biotechnology Inc. (the "Licensor"). By installing or otherwise using CryoSPARC (the "Software"), you agree to be bound by the terms and conditions of this Agreement as may be revised from time to time at Licensor's sole discretion. If you do not agree to the terms and conditions of this Agreement, do not install or use the Software.

1. NON-COMMERCIAL USE. Licensor hereby grants to Licensee one (1) non-exclusive, non-transferable license (the "License") to install and use the Software for Non-Profit Academic Research and processing of cryo-electron microscopy ("cryo-EM") data (the "Intended Purpose").\
   "Software" includes the executable computer programs, code and any related printed, electronic and online documentation, manuals, training aids, user guides, system administration documentation and any other files that may accompany the code.

   "Non-Profit Academic Research" means practicing, making, using, improving upon, importing and exporting (but not selling, leasing or otherwise monetizing) academic or scholarly research, for individual (personal) or academic institutional research purposes, in good faith, and expressly excludes, without limitation, purposes that are intended to (or result in, whether by intent or otherwise): (i) create a commercial advantage for any Person; (ii) generate monetary compensation for products or services; (iii) generate commercialization rights for any Person; (iv) be used in an ongoing business concern; or (v) result in an ongoing business concern obtaining any intellectual property rights in any research or results linked to the Non-profit Academic Research.<br>
2. RESTRICTIONS. Licensee may not: (i) modify, enhance, reverse-engineer, decompile, disassemble or create derived forms of the Software; (ii) copy the Software; (iii) sell, sub-license, lease, assign, transmit, distribute or otherwise transfer rights in/to the Software; (iv) allow third-party use of Licensee's installation of the Software; or (v) pledge, hypothecate, alienate or otherwise encumber the Software to any third party. Use of the Software is restricted to the Intended Purpose only.<br>
3. NO WARRANTY. The Software is provided "as is" without warranty of any kind. Licensor makes no representations, warranties or covenants to Licensee, either express or implied, with respect to the Software or with respect to any Confidential Information (as defined herein) disclosed to Licensee. Licensor specifically disclaims any implied warranty or condition of non-infringement, merchantable quality or fitness for a particular purpose. Licensee acknowledges that the Software is of an experimental nature, that no particular results can be guaranteed, and that it has been advised by Licensor to undertake its own due diligence with respect to all matters arising from the Agreement.<br>
4. LIMITATION OF LIABILITY. In no event is Licensor liable for any damages on any basis, in contract, tort or otherwise, of any kind and nature whatsoever, arising in respect of this Agreement, howsoever caused, including damages of any kind and nature caused by Licensor’s negligence or by a fundamental breach of contract or any other breach of duty whatsoever. Licensee is advised to safeguard important data, use caution and not rely in any way on the correct functioning or performance of the Software and/or accompanying materials.<br>
5. NO IMPROVEMENTS. Licensor is under no obligation to provide Improvements to the Software. "Improvements" means any improvements, updates, variations, modifications, alterations, additions, error corrections, enhancements, functional changes or other changes to the Software, including, without limitation: (i) improvements or upgrades to improve software efficiency and maintainability; (ii) improvements or upgrades to improve operational integrity and efficiency; (iii) changes or modifications to correct errors; and (iv) additional licensed computer programs to otherwise update the Software.<br>
6. NO FUTURE ENTITLEMENT. Nothing in this Agreement shall be construed as creating any obligation on Licensor to continue to develop, commercialize, offer, make available or support (i) the Software; or (ii) any feature, functionality or Improvement as may be encompassed in the Software from time to time.<br>
7. SUSPENSION. Licensor reserves the right to modify, suspend or discontinue, temporarily or permanently, the Software, with or without notice and without liability to Licensee.<br>
8. OWNERSHIP. Licensor retains title to and ownership of the Software and any Improvements. Nothing in this Agreement shall be construed as granting any express or implied ownership rights to Licensee in respect of the Software, associated documentation or Confidential Information, including but not limited to any patent, copyright, trademark or other intellectual property right.<br>
9. INTELLECTUAL PROPERTY. All Intellectual Property, Intellectual Property Rights and distribution rights associated with or arising from the Software or Licensor’s Confidential Information remain exclusively with Licensor. “Intellectual Property” includes, without limitation, all technical data, designs, specifications, software, data, drawings, plans, reports, patterns, models, prototypes, demonstration units, practices, inventions, methods and related technology, processes or other information, and all rights therein, including, without limitation, patents, copyrights, industrial designs, trade-marks and any registrations or applications for the same and all other rights of intellectual property therein, including any rights for which arise from the above items being treated by the Parties as trade secrets or confidential information (the rights being “Intellectual Property Rights”).<br>
10. CONFIDENTIAL INFORMATION. “Confidential Information” means any and all confidential or proprietary information of Licensor or Licensee which may be exchanged between the Parties at any time prior to and during the term of this Agreement, including, without limitation, business and marketing information, technology, know-how, ideas, reports, techniques, methods, processes, uses, composites, skills, and configurations of any kind. Without limiting the generality of the foregoing, Licensor’s Confidential Information includes: (i) the Software, including its features, functionality, performance, application and use; (ii) the computer code underlying the Software, including source and compiled code and all associated documentation and files; (iii) information relating to the performance or quality of the Software; (iv) the details of any technical assistance provided to Licensee during the term of this Agreement; (v) any other products or service made available to Licensee by Licensor during the term of this Agreement; and (vi) information regarding Licensor’s business operations or research and development activities.\
    Neither party shall: (i) disclose, either directly or indirectly, any Confidential Information or any part thereof belonging to the other party, to any person except as is specifically contemplated in this Agreement; or (ii) use any Confidential Information or any part thereof belonging to the other party, for any purpose except as is specifically contemplated in this Agreement. The obligations of confidentiality set forth herein shall not apply to the extent that the information: (i) was already known to the relevant party without restriction at the time the information was disclosed to such party; (ii) was generally available to the public or otherwise was part of the public domain at the time of its disclosure to the relevant party; (iii) became generally available to the public or otherwise part of the public domain after its disclosure to the relevant party through no act or omission of such party; or (iv) was disclosed to the relevant party without restriction by a third party who, to the best of such party's knowledge and belief, had no obligation not to disclose such information.<br>
11. FEEDBACK. Licensee may communicate to Licensor, whether or not at Licensor’s request, suggestions and comments regarding the Software, including without limitation, performance, user interface, experiment results, and errors (collectively, “Feedback”). Licensor shall have worldwide, non-exclusive, perpetual, irrevocable, royalty-free, fully-paid up rights to use such Feedback. Without limiting the generality of the foregoing, Licensor shall have the unencumbered right to make, use, copy, modify, sell, distribute, sub-license, and create derivative works of/incorporating the Feedback as part of any product, technology, service, specification or other documentation and to publicly perform or display, import, broadcast, transmit, distribute, license, offer to sell, and sell, rent, lease or lend anonymized copies of the Feedback (and derivative works thereof) as part of any product.<br>
12. PERFORMANCE DATA AND ANALYTICS. Licensor may collect usage and performance data relating to Licensee’s installation of the Software. including, without limitation, data relating to: (i) software use, including the number of users, projects and experiments associated with an installation; (ii) error information, including error messages and user-submitted feedback; (iii) performance data, including experiment run times and failed experiments; (iv) hardware utilization, including the number of active nodes and memory usage; and (v) license status information, including confirmation of valid license status.<br>
13. TERMINATION. Licensor reserves the right to terminate this Agreement immediately and without notice in the event Licensee fails to comply with any provision of this Agreement. On termination of this Agreement, whether by reason of expiry or otherwise, Licensee shall promptly discontinue use of the Software, destroy its installation of the Software and, at Licensor's request, return the Software to Licensor at no cost to Licensor. Licensor may exercise any or or more of the remedies available to it under the terms of this Agreement, in addition to any remedy available at law. Failure of Licensor to enforce a right under this Agreement shall not act as a waiver of that right.<br>
14. GOVERNING LAW. This Agreement is made in Ontario and governed by and construed in accordance with the laws of the Province of Ontario and the federal laws of Canada applicable therein. The Parties attorn to the exclusive jurisdiction of the Courts of the Province of Ontario.<br>
15. SURVIVAL. The provisions of subsections 2-12 shall survive termination of this Agreement.<br>
16. SEVERABILITY. If any term, covenant, condition or provision of this Agreement is held by a court of competent jurisdiction to be invalid, void or unenforceable, it is the Parties' intent that such provision be reduced in scope by the court only to the extent deemed necessary by that court to render the provision reasonable and enforceable and the remainder of the provisions of this Agreement will in no way be affected, impaired or invalidated as a result.<br>
17. NO AGENCY. No provision of this Agreement or action by the Parties will establish or be deemed to establish any partnership, joint venture, principal-agent or employer-employee relationship in any way, or for any purpose, between Licensor and Licensee.<br>
18. ENTIRE AGREEMENT. This Agreement including all schedules hereto, constitutes the entire agreement between the Parties concerning the subject matter hereof and supersedes all prior or collateral agreements, communications, representations, understandings, negotiations and discussions, oral or written.

© 2020 Structura Biotechnology Inc.


# CryoSPARC Architecture and System Requirements

Description of CryoSPARC HPC software system architecture, typical setups (e.g., workstation, cluster).

## CryoSPARC System Architecture Overview

CryoSPARC is a backend and frontend high-performance computing software system that provides data processing and image analysis capabilities for single particle cryo-EM, along with a rich browser-based user interface and command line tools.

CryoSPARC can be deployed on-premises or in the cloud.

{% hint style="warning" %}
CryoSPARC is designed to be run only within a trusted private network. CryoSPARC instances are not security-hardened against malicious actors on the network and should never be hosted directly on the internet or an untrusted network without a separate controlled authentication layer.
{% endhint %}

### Master-worker pattern

The system is based on a master-worker pattern.

* The master processes (web application, core application and MongoDB database) run together on one machine (**master** node). The master node requires relatively lightweight resources (4+ CPUs, 16GB+ RAM, 250GB+ HDD storage)
* Worker processes run on any available/configured machine that has NVIDIA GPUs (**worker** node). The worker is responsible for all actual computation and data handling and is dispatched by the master node.

{% hint style="info" %}
The same node can function as both master and worker.
{% endhint %}

The master-worker architecture allows CryoSPARC to be installed and scaled up flexibly on a variety of hardware, including a single workstation, groups of workstations, cluster nodes, HPC clusters, cloud nodes, and more.

![Core components included in the CryoSPARC system](/files/-M7DHJA1-XK7uXA6RxHc)

## Typical CryoSPARC System Setups

{% hint style="info" %}
CryoSPARC can support a heterogeneous mixture of all typical setups in a single instance. This means you can start with installing CryoSPARC on a single workstation, then connect a worker node or cluster as your data processing requirements scale.
{% endhint %}

### Single Workstation

![Single CryoSPARC workstation example, where the master and worker processes run on a single machine.](/files/-M7DHJA22tU6Tn3Ezizv)

Both the CryoSPARC **master** and CryoSPARC **worker** processes may run on the same machine. The only requirement is that GPU resources are available for the CryoSPARC worker processes. This is the simplest setup.

### Master-Worker

![Single-master, multiple-worker example. All nodes must have access to a shared file system. Nodes may be assigned to a dedicated lane (such as lanes B and C in this example) or combined into a common lane (lane A).](/files/cmcHWSPrYpveSkU8IdR7)

In the master-worker setup, the worker processes run on one or more GPU servers ("nodes"). The master processes may be run on a light-weight dedicated server or on a GPU server that is configured similarly to a single workstation (see [above](#single-workstation)). This is the most flexible setup for installing CryoSPARC. There are three main requirements for this setup, which are also explained in greater detail in the [installation sections of this document](/setup-configuration-and-management/cryosparc-installation-prerequisites):

{% hint style="danger" %}
**1) All nodes have access to a shared file system.** This file system is where the project directories are located, allowing all nodes to read and write intermediate results as jobs start and complete.

**2) The master node** has password-less SSH access to each of the worker nodes. SSH is used to execute jobs on the worker nodes from the master node.

**3) All worker nodes** have TCP access to 10 consecutive ports on the master node. These ports are used for metadata communication via HTTP API requests.
{% endhint %}

{% hint style="info" %}
If master processes are run on a GPU server, heavy GPU processing loads can lead to instability if the GPU worker node hangs or runs out of RAM, causing the master processes running the web application and database to also hang.
{% endhint %}

In this *master-worker* pattern

* processing jobs are queued to a specific scheduler *lane*
* a scheduler *lane* is a collection of one or more (GPU) worker nodes
* each worker node is associated with no more than one lane

One may combine multiple worker nodes into a lane if there is no concern over which specific node will process a given CryoSPARC job.

One may prefer single-node lanes if more direct control is desired over where a given job will run. This may be the case when hardware specifications, such as the amount of RAM or GPU models, vary significantly between nodes. A potential drawback is that a job queued to such a lane would remain queued to the selected lane even if resources are idle on another lane.

### Clusters

![CryoSPARC cluster integration example where both nodes have access to a shared file system](/files/-M7DHJA4o5Uj7cW4cRcB)

The master node can also spawn or submit jobs to a cluster scheduler system (e.g., [Slurm Workload Manager](https://slurm.schedmd.com/overview.html)). This integration is transparent, and works similar to the master-worker setup explained above, except all resource scheduling is handled by the cluster scheduler, and CryoSPARC's scheduler is only used for orchestration and management of jobs. Similar requirements are present:

{% hint style="danger" %}
**1) All nodes** have access to a shared file system. This file system is where the project directories are located, allowing all nodes to read and write results as jobs start and complete.

**2) All worker nodes** have TCP access to 10 consecutive ports on the master node (default ports are 39000-39009). These ports are used for metadata communication via HTTP Remote Procedure Call (RPC) based API requests.
{% endhint %}

{% hint style="info" %}
For a **cluster** setup, the master node can be a regular cluster node (or even a login node) if this makes networking requirements easier, but the CryoSPARC master processes must be run continuously. If the master is to be run on a regular cluster node, the node may need to be requested from your scheduler in interactive mode or for an indefinitely running job.
{% endhint %}

{% hint style="info" %}
Project directories are created in locations specified by CryoSPARC users. If administering a multi-user cluster instance, ensure that users create project directories in locations where both the master and worker nodes have access.
{% endhint %}

#### Supported cluster schedulers

CryoSPARC supports most cluster schedulers, including SLURM, SGE and PBS. Please see [here for more details about how CryoSPARC connects with a cluster system](/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc#connect-a-cluster-to-cryosparc).

## CryoSPARC System Requirements

The following are requirements for every master and worker node in the system unless otherwise specified.

|        Component | Requirement                                                                                                                                                                                                       |
| ---------------: | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|     Architecture | x86-64 (Intel or AMD)                                                                                                                                                                                             |
| Operating System | Modern Linux OS. Please see the section on [Operating System](https://guide.cryosparc.com/setup-configuration-and-management/hardware-and-system-requirements#operating-system) for more details and limitations. |
|            Shell | [Bash](https://www.gnu.org/software/bash/)                                                                                                                                                                        |
|     User Account | `cryosparcuser`                                                                                                                                                                                                   |
|         Software | Nvidia driver (worker nodes only). [See details on Nvidia and CUDA requirements](https://guide.cryosparc.com/setup-configuration-and-management/pages/-M7DHIJrIWYpsjbmcVFX#1.-nvidia-driver).                     |
|       Filesystem | Shared file system across all nodes                                                                                                                                                                               |

### CryoSPARC Master Node Requirements

The following are requirements specific to the master node.

|              Component | Minimum Requirement           | Recommended                    |
| ---------------------: | ----------------------------- | ------------------------------ |
|                **CPU** | 4+ cores                      | 8+ cores at 2.8GHz+            |
|                **RAM** | 16GB+                         | 32GB DDR4                      |
|     **System Storage** | 250GB+ HDD                    | 500GB SSD                      |
| **Fast Local Storage** | Not Required                  | Not Required                   |
|                **GPU** | Not Required                  | Not Required                   |
|            **Network** | 1Gbps link to storage servers | 10Gbps link to storage servers |

A 10Gbps connection is recommended to the storage servers given raw cryo-EM movies can be several TB in size, and I/O bottlenecks are more of a concern than processing power for pre-processing jobs in CryoSPARC.

Although a CPU with a higher core count is recommended, a CPU with a faster clock rate is more advantageous due to how master processes are implemented.

Enough System Storage is required to host the `cryosparc_master` installation package and database folder. Each CryoSPARC project occupies between 100MB and 5GB of database storage, depending on the size of the project. 500GB is enough for approximately 200 medium-sized projects. *Note that this excludes the space required for CryoSPARC project data in bulk storage, which could be in terabytes for larger projects.*

### Worker Node/Cluster Worker Minimum Requirements

The following are requirements for each worker node/cluster worker.

|                Component | Minimum Requirement                                                                      | Recommended                                   |
| -----------------------: | ---------------------------------------------------------------------------------------- | --------------------------------------------- |
|                  **CPU** | 2+ cores **per GPU**                                                                     | 4 cores **per GPU**                           |
| **CPU Memory Bandwidth** | 50+ GB/s                                                                                 | 100+ GB/s                                     |
|                  **RAM** | 32GB+ **per GPU**                                                                        | 64GB DDR4 **per GPU**                         |
|       **System Storage** | 25GB+ HDD                                                                                | 50GB+ SSD                                     |
|   **Fast Local Storage** | 1TB SSD                                                                                  | 2TB PCIe SSD                                  |
|                  **GPU** | 1+ NVIDIA GPU with [CC 3.5+](https://developer.nvidia.com/cuda-gpus#compute), 11GB+ VRAM | 1+ NVIDIA Tesla V100, RTX2080Ti, RTX3090, etc |
|              **Network** | 1Gbps link to storage servers                                                            | 10Gbps link to storage servers                |

High CPU memory bandwidth is especially important for [CryoSPARC Live preprocessing](/live/prerequisites-and-compute-resources-setup#preprocessing-lane).

System RAM is very important for worker nodes and should scale proportionately to the number of GPUs available for processing on the system.

Enough System Storage is required to host the `cryosparc_worker` installation package.

Fast local storage is also necessary as reconstruction jobs require random access to particle images. SSDs provide high throughput in this context. See the section on [Solid State Storage](https://guide.cryosparc.com/setup-configuration-and-management/hardware-and-system-requirements#solid-state-storage-ssds) for more details.

### Operating System

For CryoSPARC v5.0+, the operating system must support GLIBC 2.28 or greater. Therefore, the oldest compatible operating systems are Rocky/RHEL 8 and Ubuntu 20.04. For Ubuntu, version 22.04 or 24.04 is recommended.

We recommend running CryoSPARC on an up-to-date long-term support Linux distribution, such as Ubuntu (22.04, 24.04), Rocky Linux (8, 9, 10) or a related distribution.

{% hint style="warning" %}
As of June 2026, CryoSPARC versions up to 5.0.6 are incompatible with the version 7 kernel that is the default for Ubuntu 26.04.
{% endhint %}

### Disks and compression

Fast disks are a necessity for processing cryo-EM data efficiently. Fast **sequential** read/write throughput is needed during **pre-processing** stages (e.g., motion correction) where the volume of data is very large (tens of TB) while the amount of computation is relatively low (sequential processing for motion correction, CTF estimation, particle picking, etc.)

Spinning disk arrays in a RAID configuration are used to store large raw data files, and often cluster file systems are used for larger systems. As a rule of thumb, to saturate a 4-GPU machine during pre-processing, a **sustained sequential read of 1000MB/s is required**.

Compression can greatly reduce the amount of data stored in movie files, and also greatly speeds up preprocessing because decompression is actually faster than reading uncompressed data straight from disk. Typically, counting-mode movie files are stored in LZW compressed TIFF format without gain correction, so that the gain reference file is stored separately and must be applied on-the-fly during processing (which is supported by CryoSPARC). Compressing gain corrected movies can often result in much **worse** compression ratios than compressing pre-gain corrected (integer count) data.

CryoSPARC supports LZW compressed TIFF format, EER format and BZ2 compressed MRC format natively. In either case, the gain reference must be supplied as an MRC file. TIFF, EER and BZ2 compression are implemented as multi-core decompression streams on-the-fly.

### Solid State Storage (SSDs)

SSD space is optional on a per-worker node basis but is **highly recommended** for worker nodes that will be running refinements and reconstructions using particle images. Nodes reserved for pre-processing (motion correction, particle picking, CTF estimation, etc) **do not** need to have an SSD.

CryoSPARC particle processing algorithms rely on random-access patterns and multiple passes through the data, rather than sequentially reading the data at once. Using a storage medium that allows for fast random reads will speed up processing dramatically.

CryoSPARC manages the SSD cache on each worker node transparently. Files are automatically cached, re-used across the same project and deleted if more space is needed. [Please see the SSD Caching guide for more information.](/setup-configuration-and-management/software-system-guides/tutorial-ssd-particle-caching-in-cryosparc)

The size of your typical single particle cryo-EM datasets will inform the size of SSD you choose to use. For a sample calculation, see:

{% embed url="<https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/tutorial-ssd-particle-caching-in-cryosparc#hardware>" %}

### Graphical Processing Units (GPUs)

At least one worker node must have GPUs available to run the complete set of CryoSPARC jobs. Non-GPU workers may run CPU-only jobs.

The GPU memory (VRAM) in each GPU limits the maximum particle box size for reconstruction. Typically, a GPU with 12GB VRAM can handle a box size of up to 700^3, and up to 1024^3 in some job types.

Please ensure each connected worker includes a recent version of the Nvidia Driver compatible with your GPU. See [this section](https://guide.cryosparc.com/setup-configuration-and-management/pages/-M7DHIJrIWYpsjbmcVFX#1.-nvidia-driver) for details. [Download the latest driver for your GPUs here](https://www.nvidia.com/Download/index.aspx). Visit [Troubleshooting](/setup-configuration-and-management/troubleshooting#gpu-issues) to resolve common GPU errors.

#### Selecting GPUs

When acquiring GPUs to use with CryoSPARC, the following considerations may be useful.

* CryoSPARC almost exclusively uses single-precision operations on the GPU. As such, consumer cards generally have a much better price/performance ratio than enterprise cards. Enterprise cards do have their own benefits such as reliability, better cooling for servers, longer support timelines, and compatibility with other applications that may use double-precision math.
* The most important metric of a GPU is the device VRAM memory bandwidth. This is the rate at which the GPU can read and write from its own memory. This is generally more important than GPU core count or clock speed, as almost all operations on the GPU are memory-bandwidth limited. When selecting GPUs, this is the primary metric to compare (along with price). For example, the NVIDIA RTX 3090 has 24GB of memory at 936 GB/s bandwidth, the NVIDIA A100 has up to 80GB memory at 1935 GB/s bandwidth, and the NVIDIA A4000 has 16GB at 448 GB/s.
* GPU memory size is the main limiting factor in terms of the box-sizes that can be handled during a 3D refinement. Other than this, memory size does not have any impact on speed. 11GB consumer cards can generally handle all processing steps (including motion correction of K3 data, etc) for particle box sizes up to 600^3.
* GPU-CPU interconnect bandwidth (eg. PCIE) is generally not a bottleneck (e.g., for most job types, we get similar benchmark performance on 8x or 16x PCIE lanes) but IO bandwidth reading data from cluster storage/local SSD is usually a significant factor in performance. This is especially true for preprocessing and CryoSPARC Live, as movies, micrographs, and particles need to be read, written, and transferred rapidly to keep up with collection and GPUs can process the data very quickly.
* In many cases, older or slower GPUs can often perform almost equally as well as the newest, fastest GPUs because most computations in CryoSPARC are **not** bottlenecked by GPU compute speed, but rather by GPU memory bandwidth and disk I/O speed.

### Browser Requirements

The CryoSPARC web interface works best on the latest version of [Google Chrome](https://www.google.com/chrome/index.html). [Firefox](https://www.mozilla.org/en-US/firefox/new/) and [Safari](https://www.apple.com/safari/) are also an option, although some features may not work as intended. Internet Explorer is not supported. [See this guide](/setup-configuration-and-management/how-to-download-install-and-configure/accessing-cryosparc) for more information on accessing the CryoSPARC web interface.

## Additional Configuration Notes

### Network Accessibility

Network security must be an important factor in the installation and ongoing management of any CryoSPARC instance. Access to the network that hosts a CryoSPARC instance must be carefully controlled, as CryoSPARC instances are not security-hardened against malicious actors on the network.

CryoSPARC is designed to be run only within a trusted private network. CryoSPARC instances should never be directly hosted on the internet or an untrusted network, without a separate controlled authentication layer.

CryoSPARC’s User Interface does include a user management system, and CryoSPARC user accounts and passwords help control access to the interface within a trusted private network, but please note that CryoSPARC passwords are not intended as a barrier against malicious access.

### Root Access

The CryoSPARC system is specifically designed not to require root access to install or use. The reason for this is to avoid security vulnerabilities that can occur when a network application (web interface, database, etc.,) is hosted as the root user. For this reason, the CryoSPARC system must be installed and run as a regular UNIX user (`cryosparcuser`), and all input and output file locations must be readable and writable as this user. In particular, this means that project input and output directories that are stored within a regular user's home directory need to be accessible by `cryosparcuser`, or else (more commonly) another location on a shared file system must be used for CryoSPARC project directories.

### Multi-user environment

If you are installing the CryoSPARC system for use by many users (for example within a lab), there are two options:

#### Using UNIX Groups

Create a new regular user (`cryosparcuser`) and install and run CryoSPARC as this user. Create a CryoSPARC project directory (on a shared file system) where project data will be stored, and create sub-directories for each lab member. If extra security is necessary, use UNIX group privileges to make each sub-directory read/writeable only by `cryosparcuser` and the appropriate lab member's UNIX account. Within the CryoSPARC command-line interface, create a CryoSPARC user account for each lab member, and have each lab member create their projects within their respective project directories. This method relies on the CryoSPARC web application for security to limit each user to see only their own projects. This is not guaranteed security, and malicious users who try hard enough will be able to modify the system to be able to see the projects and results of other users.

#### Using Separate CryoSPARC Instances

If each user must be guaranteed complete isolation and security of their projects, each user must install CryoSPARC independently within their own home directories. Projects can be kept private within user home directories as well, using UNIX permissions. Multiple single-user CryoSPARC master processes can be run on the same master node, and they can all submit jobs to the same cluster scheduler system. This method relies on the UNIX system for security and is more tedious to manage but provides stronger access restrictions. Each user will need to have their own CryoSPARC license ID in this case.

### Deploying CryoSPARC on AWS

CryoSPARC can be deployed on-premises or in the cloud. See below for a guide on deploying CryoSPARC on AWS resources.

{% content-ref url="/pages/-M\_L\_k\_-56oUni\_efXg\_" %}
[Deploying CryoSPARC on AWS](/setup-configuration-and-management/cryosparc-on-aws)
{% endcontent-ref %}

### Database and Command API Security

{% hint style="warning" %}
CryoSPARC is designed to run within a trusted private network. CryoSPARC instances are not security-hardened against malicious actors on the network and should never be exposed to the Internet or hosted on an untrusted network.

The information in this section, Database and Command API Security, applies to CryoSPARC v4.0+.
{% endhint %}

CryoSPARC v4.0 introduces additional authentication to reduce the likelihood of accidental mis-use by actors on a large institution/multi-user shared network.

#### Database Security

CryoSPARC's MongoDB database runs with access control enabled: Requests to read or write from the database must be authenticated with a username and password. CryoSPARC sets this password automatically and uses it for internal master → master and worker → master requests. Note that communication between the database and other CryoSPARC services is not encrypted.

Enable or disable MongoDB access control by setting `CRYOSPARC_DB_ENABLE_AUTH` variable in `cryosparc_master/config.sh` and `cryosparc_worker/config.sh` to `true`(default) or `false`.

To connect to the database with access control, use `cryosparcm mongo` to access the Mongo shell, or use `cryosparcm icli` to access an interactive Python client.

#### Command API Security

CryoSPARC's command API[^1] server (executes actions triggered by the web application including creating and modifying projects and jobs) also requires authentication.

The command server expects the CryoSPARC License ID in the`License-ID` header of incoming web requests. CryoSPARC includes this automatically in internal requests to the API. Requests with missing or incorrect license ID will be rejected. Note that communication between the command server and other CryoSPARC services is not encrypted.

You provide the License ID during CryoSPARC master and worker package installation. The license is written in plain-text to `config.sh` in the installation directories. The license ID in the worker installation must match the master installation.

#### Optional password argument

CryoSPARC allows creating users and updating users from the command line. Rather than specifying a `--password` flag during these operations, you may omit it to instead run a secure password prompt.

## Example Systems

We **do not** currently partner with any specific hardware vendors to sell machines with CryoSPARC pre-installed.

### Example Hardware Systems

Below are details of example workstations that meet or exceed the minimum requirements specified above, including those we use internally for development and testing.

{% tabs %}
{% tab title="Example 4-GPU" %}

|    Component | Hardware Product                                                               |
| -----------: | ------------------------------------------------------------------------------ |
|          CPU | 32 Cores (base clock 2.8GHz+), e.g, AMD Threadripper 3975X                     |
|       Memory | 256GB DDR4 @ 3200MHz                                                           |
|      Storage | 4TB PCIe SSD (cache); 200TB RAID 6 storage server via 10Gbps link (raw movies) |
|          GPU | 4x NVIDIA Quadro GV100, or 4x NVIDIA Tesla V100 or 4x NVIDIA RTX 8000          |
| {% endtab %} |                                                                                |

{% tab title="Example 2-GPU" %}

|     Component | Hardware Product                                                            |
| ------------: | --------------------------------------------------------------------------- |
|           CPU | 16 Cores (base clock 3.0GHz+)                                               |
|        Memory | 128GB DDR4                                                                  |
|       Storage | 2TB PCIe SSD (cache); HDD storage server in RAID configuration (raw movies) |
|           GPU | 2x NVIDIA RTX 3090                                                          |
|  {% endtab %} |                                                                             |
| {% endtabs %} |                                                                             |

[^1]: Application Programming Interface


# CryoSPARC Installation Prerequisites

Before installing CryoSPARC, ensure these six requirements are met.

## 1. Nvidia Driver

CryoSPARC worker installations on workstations, dedicated GPU nodes or clusters require a recent version of the Nvidia driver and a Nvidia GPU. The list below specified the required Nvidia Driver version for a range of CryoSPARC versions.

* **CryoSPARC v5.0+**
  * Requires **Nvidia Driver version 570.26 or newer**. Note that NVIDIA Blackwell devices are only compatible with the open driver.
  * A system CUDA installation is not needed. CryoSPARC includes CUDA 12.8 which drops support for NVIDIA GPUs with compute capability 3.5 (Kepler). Only GPUs with compute capability 5.0 (Maxwell) to 12.0 (Blackwell) are supported.
* **CryoSPARC v4.4 to CryoSPARC v4.7**
  * Requires **Nvidia Driver version 520.61.05 or newer.**
  * A system CUDA installation is not needed. CryoSPARC includes CUDA 11.8.0.
* **CryoSPARC \<v4.4**
  * Requires a system CUDA installation. CryoSPARC runs with CUDA version 11, and we recommend toolkit version 11.8 and the corresponding Nvidia Driver version.

Please follow instructions specific to the CryoSPARC worker node's Linux distribution to install the Nvidia driver. Visit [Troubleshooting](/setup-configuration-and-management/troubleshooting#gpu-issues) for common GPU errors.

## 2. Common Unix User Account

The same CryoSPARC-associated, non-privileged Linux account must be available on the CryoSPARC master and all worker nodes.

{% hint style="danger" %}
Do not use the root account to install, update or manage a CryoSPARC instance.
{% endhint %}

{% hint style="warning" %}
Execute all command line instance management tasks, such as updates or startup, under the Unix account that runs the CryoSPARC instance. Failure to do so may render the CryoSPARC instance inoperative.
{% endhint %}

{% hint style="info" %}
You don't need to have a dedicated Unix user (e.g., `cryosparcuser`), to run and install CryoSPARC -- you can use your own Linux account, but do **not** use the root account. Using your own Linux account makes sense when you are installing CryoSPARC for yourself, and you don't plan on having any other users use the same instance.
{% endhint %}

{% hint style="info" %}
The CryoSPARC-associated Linux account must be associated with the same numeric UID on all nodes.
{% endhint %}

In a master-worker setup, the CryoSPARC master node will use SSH to access the worker node and execute a bash script that will run the job a user has queued to that machine. Some lightweight job types queue directly to the master node, in which case the CryoSPARC master process will execute the job using a Python subprocess. If a user queues a job to a cluster, the CryoSPARC master process will submit a cluster job via the cluster workload scheduler's job submission system (for example via the `sbatch` command on a SLURM cluster).

For the purposes of this documentation, `cryosparcuser` represents the Linux account that owns the CryoSPARC processes.

## 3. Password-less SSH Access

‌Set up SSH access between the master node and each standalone worker node. The `cryosparcuser` account should be able to SSH without a password (using a SSH key-pair) into all non-cluster worker nodes.

### Setting up password-less SSH access to a remote workstation

Set up SSH keys for password-less access (only if you currently need to enter your password each time you ssh into the compute node).

If you do not already have SSH keys generated on your local machine, use `ssh-keygen` to do so.

#### Open a terminal prompt on your local machine, and enter:

```
ssh-keygen -t rsa -N "" -f $HOME/.ssh/id_rsa
```

{% hint style="info" %}
This will create an RSA key-pair with no passphrase in the default location.‌
{% endhint %}

#### Copy the RSA public key to the remote compute node for password-less login:

```
ssh-copy-id remote_username@remote_hostname
```

{% hint style="info" %}
*`remote_username` and `remote_hostname` are your username and the hostname that you use to SSH into your compute node. This step will ask for your password.*
{% endhint %}

## 4. Open TCP Ports

The port range is configurable during install time. Select a suitable range of ten consecutive network ports on the CryoSPARC master computer that

1. does **not** coincide with ports used by non-CryoSPARC services running on this computer.
2. does **not** overlap with the port range of another CryoSPARC instance that may be running on this computer.
3. does **not** overlap with the computer's "ephemeral" port range. A Linux computer's ephemeral port range can be displayed with the command

   `cat /proc/sys/net/ipv4/ip_local_port_range`

The following table details the purpose of each port, assuming a `--port 61000` installation parameter.

| 61000 | CryoSPARC web application                                                                            |
| ----- | ---------------------------------------------------------------------------------------------------- |
| 61001 | MongoDB database; needs be accessible from CryoSPARC workers                                         |
| 61002 | REST[^1] API[^2] web server; in CryoSPARC v4 or older, needs to be accessible from CryoSPARC workers |
| 61003 | Command Visualization (Vis) server                                                                   |
| 61004 | Redis cache and coordination service; needs to be accessible from CryoSPARC workers                  |
| 61005 | Supervisor service management server                                                                 |
| 61006 | CryoSPARC web application API[^2] server                                                             |
| 61007 | Reserved *(Not Used)*                                                                                |
| 61008 | Reserved *(Not Used)*                                                                                |
| 61009 | Reserved *(Not Used)*                                                                                |
|       |                                                                                                      |

### Checking ports in use

To see whether certain ports are being used on your master node, run a command like

`sudo ss -anp | grep ":6100[0-9][^0-9]" | sed 's/\s+/ /g'`

where `ss` output is filtered with `grep` for a port number pattern.

### Testing open ports

To test if a TCP port is open (for example, to test if there is a firewall blocking the port), run a `telnet` command from another computer inside the network. If you see any response other than the one below (e.g., a timeout or a denial), it may indicate that the port is not listening or is blocked.

```
$ telnet cryosparc.server 61000
Trying 192.168.64.49...
Connected to cryoem5.slush.sandbox.
Escape character is '^]'.
$ ^C
Connection closed by foreign host.
```

## 5. Shared File System

The major requirement for installation is that all nodes (including the master) be able to access the same shared file system(s) at the same absolute path. These file systems (typically cluster file systems or NFS mounts) will be used for loading input raw data into jobs running on various nodes, as well as saving output data from jobs into projects.

![Example of a Master-Worker setup where all nodes have access to the same shared filesystem](/files/-M7DHJA3AfcYz1p0da27)

{% hint style="info" %}
Each project created by a user is associated with a single project directory that all CryoSPARC nodes must be able to read from and write to. All users should create project directories in locations where both the master and worker nodes have access.
{% endhint %}

{% hint style="info" %}
CryoSPARC project directories need to be stored on a filesystem that supports symbolic links.
{% endhint %}

## 6. Outbound HTTPS Internet Access

CryoSPARC requires internet access from the main process to verify your license and perform updates. At minimum, CryoSPARC should have access to our license server at [`https://get.cryosparc.com/`](https://get.cryosparc.com/). See [here](https://guide.cryosparc.com/setup-configuration-and-management/how-to-download-install-and-configure/accessing-cryosparc#appendix-d-custom-ssl-certificate-authority-bundle) for more details.

[^1]: Representational State Transfer, a standard application server design format

[^2]: Application Programming Interface


# How to Download, Install and Configure

Meeting system requirements, obtaining a License ID, and downloading & installing CryoSPARC.

## Step 1: Confirm Prerequisites

Review

{% content-ref url="/pages/-M7DHIJqfv6KLfsg4OXA" %}
[CryoSPARC Architecture and System Requirements](/setup-configuration-and-management/hardware-and-system-requirements)
{% endcontent-ref %}

and complete all applicable

{% content-ref url="/pages/-M7DHIJrIWYpsjbmcVFX" %}
[CryoSPARC Installation Prerequisites](/setup-configuration-and-management/cryosparc-installation-prerequisites)
{% endcontent-ref %}

## Step 2: Obtain a CryoSPARC License ID

CryoSPARC is available free of charge for [non-profit academic use](/licensing). To download the software, you will need a CryoSPARC License ID, which can be requested via the form at <https://cryosparc.com/download/>.

Please contact <sales@structura.bio> for inquiries relating to for-profit or industry licensing, including academic-industry collaborations on proprietary projects and fee-for-service data processing.

{% content-ref url="/pages/-M7DHIJtPASw90XL77Fg" %}
[Obtaining A License ID](/setup-configuration-and-management/how-to-download-install-and-configure/obtaining-a-license-id)
{% endcontent-ref %}

## Step 3: Download and Install CryoSPARC

Please see the detailed installation steps applicable to your system setup.

{% content-ref url="/pages/-M7DHIJu6M1ubOlnm8Vp" %}
[Downloading and Installing CryoSPARC](/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc)
{% endcontent-ref %}


# Obtaining A License ID

Fill out the form to obtain a License ID required for installing and using CryoSPARC.

To obtain a License ID for CryoSPARC, go to [cryosparc.com/download](http://cryosparc.com/download), fill out the form and submit it.

![License request form](/files/-M7DHJkpN9um4h8ffqDp)

![Confirmation message once form is submitted](/files/-M7DHJkqRDj2zQFwqdwP)

Once you fill out the form, you will receive an email confirming your request.

![Email notifying that your request was successful](/files/-M7DHJkrZJnUfvLsh99t)

We endeavour to respond to all requests within 24 business hours. Once verified, you will receive a second email containing your License ID.

![Email containing your cryoSPARC License ID and next steps](/files/-M7DHJksfIUsyV8bSald)


# Downloading and Installing CryoSPARC

Downloading and installing the cryosparc\_master and cryosparc\_worker packages.

## Overview

Once you've reviewed the CryoSPARC Installation Prerequisites and have decided on an architecture suitable for your situation, the installation process can be distilled to a few main steps which are further detailed below:

1. Download the `cryosparc_master` package
2. Download the `cryosparc_worker` package
3. Install the `cryosparc_master` package on the master node
4. Start cryoSPARC for the first time
5. Create the first administrator user
6. \[Optional but recommended] Set up recurring backup of the CryoSPARC database
7. **If you chose to install CryoSPARC via the single-workstation method, at this point, you are finished installing CryoSPARC. Otherwise, continue.**
8. Log onto a worker node and install the `cryosparc_worker` package (installing the `cryosparc_worker` package requires an Nvidia GPU and Nvidia Driver 520.61.05 or newer)
9. Depending on whether you're installing CryoSPARC on a cluster or on a standalone worker machine, connect the worker node to CryoSPARC via `cryosparcm cluster connect` or `cryosparcw connect`
10. **CryoSPARC is now fully installed.** You may also choose to add more standalone worker machines at any time by using the `cryosparcw connect` utility.

{% hint style="warning" %}
You will need at least 15 GB of space to download and install the `cryosparc_master` and `cryosparc_worker`packages
{% endhint %}

## Prepare for Installation

### **Log into the workstation** where you would like to install and run CryoSPARC

```
ssh <cryosparcuser>@<cryosparc_server>
```

{% hint style="info" %}
`<cryosparcuser> = cryosparcuser`

* The username of the account that cryoSPARC is to be installed by.

`<cryosparc_server> = uoft`

* The hostname of the server where cryoSPARC will be installed on.
  {% endhint %}

### Select a suitable directory for installation

`cryosparc_master/`and/or `cryosparc_worker/`directories will be created inside the selected directory during a subsequent step. Select the directory carefully because the `cryosparc_master/`and `cryosparc_worker/` directories cannot be moved following installation.

The absolute path to the installation directory must not exceed 83 characters (after dereferencing sybolic links in the path, if applicable). Suppose one selected and, if necessary, created the `/sw/cryosparc/`directory for this purpose:

```bash
# enter the selected directory
cd /sw/cryosparc
# confirm the max path length is not exceeded
echo -n $(pwd -P) | wc -c
```

## **Download and Extract the `cryosparc_master` and `cryosparc_worker` Packages** Into Your Installation Directory

### **Define your License ID** as an environment variable:

```bash
LICENSE_ID="your-unique-license-id"
```

{% hint style="info" %}
Your unique CryoSPARC license ID

* should be included in an e-mail we sent to you
* should look similar to the example `682437fb-d6ae-47b8-870b-b530c587da94`
* can be used for the installation and operation of a single CryoSPARC master instance and all CryoSPARC workers associated with that single master instance
  {% endhint %}

{% hint style="warning" %}
CryoSPARC download or installation commands that include or refer to `$LICENSE_ID` will fail if the `LICENSE_ID` variable is not defined in the current shell's environment.
{% endhint %}

### Use `curl` to **download the two files** into tarball archives

You can download the latest public CryoSPARC version, or you can specify a particular version to install. For the latest version, use the below commands. To specify a version, replace `VERSION="latest"` with `VERSION="vX.Y.Z"` (for example, `VERSION="v4.7.1"` ). See the [changelog ](https://cryosparc.com/updates)for all available versions.

```
VERSION="latest"
curl -L https://get.cryosparc.com/download/master-$VERSION/$LICENSE_ID -o cryosparc_master.tar.gz
curl -L https://get.cryosparc.com/download/worker-$VERSION/$LICENSE_ID -o cryosparc_worker.tar.gz
```

### **Extract the downloaded files**:

```
tar -xf cryosparc_master.tar.gz cryosparc_master
tar -xf cryosparc_worker.tar.gz cryosparc_worker
```

**Note:** After extracting the worker package, you may see a second folder called `cryosparc2_worker` (note the `2`) containing a single `version` file. This is here for backward compatibility when upgrading from older versions of CryoSPARC and is not applicable for new installations.

You may safely delete the `cryosparc2_worker` directory.

## **Install The `cryosparc_master` Package**

Follow these instructions to install the master package on the master node. If you are following the Single Workstation setup, once you run the installation command, CryoSPARC will be ready to use.

{% hint style="danger" %}
After installation and startup, the `cryosparc_master` software [exposes a number of network ports](https://guide.cryosparc.com/setup-configuration-and-management/how-to-download-install-and-configure/pages/-M7DHIJrIWYpsjbmcVFX#4.-open-tcp-ports) for connections from other computers, such as CryoSPARC worker nodes and computers from which users browse the CryoSPARC user interface. You must ensure that these ports cannot be accessed *directly* from the internet. The CryoSPARC guide contains [suggestions](/setup-configuration-and-management/how-to-download-install-and-configure/accessing-cryosparc) on implementing access to the user interface.
{% endhint %}

{% hint style="warning" %}
If you are installing CryoSPARC on a single workstation, you can use the **Single Workstation** instructions below that simplify installation to a single command. Otherwise, use the **Master Node Only** instructions here, and then continue with the further steps to install on worker nodes separately.
{% endhint %}

{% hint style="warning" %}
Select and monitor the filesystem that hosts the `$CRYOSPARC_DB_PATH` directory carefully. The CryoSPARC instance will fail if the filesystem is allowed to fill up completely. Recovery from such a failure may take significant time for large instance. The directory can be specified during installation with the "--`dbpath"` parameter of the `install.sh` command. Otherwise, the default is a directory named `cryosparc_database` inside the directory where the software tar packages were unpacked.

The CryoSPARC database holds metadata, images, plots and logs of jobs that have run. If the database becomes corrupt or lost due to to user error or filesystem issues, the instance must be recovered. A CryoSPARC instance can be recovered as follows:

1. For CryoSPARC versions up to v4.7: In a new/fresh CryoSPARC instance, import the **project directories** of projects from the original instance. This should resurrect all jobs, outputs, inputs, metadata, etc. **Please note:** user and instance level settings and configuration will be lost. Users, worker lanes, and scheduler configurations will need to be re-created in the new instance.
2. For CryoSPARC v5.0+: The new `cryosparcm recover` command can be used to automatically recover an instance if the database is lost. See [Instance Recovery](/setup-configuration-and-management/software-system-guides/guide-instance-recovery-v5.0) for more details.
   {% endhint %}

{% tabs %}
{% tab title="Single Workstation (Master and Worker combined)" %}

### **Single Workstation** CryoSPARC Installation

```
cd cryosparc_master

./install.sh    --standalone \ 
                --license $LICENSE_ID \ 
                --worker_path <worker path> \ 
                --ssdpath <ssd path> \ 
                --initial_email <user email> \
                --initial_username "<login username>" \
                --initial_firstname "<given name>" \
                --initial_lastname "<surname>" \
                [--port <port_number>] \
                [--initial_password <user password>]
```

**Example command execution**

```
./install.sh    --standalone \
                --license $LICENSE_ID \
                --worker_path /u/cryosparc_user/cryosparc/cryosparc_worker \
                --ssdpath /scratch/cryosparc_cache \
                --initial_email "someone@structura.bio" \
                --initial_password "Password123" \
                --initial_username "username" \
                --initial_firstname "FirstName" \
                --initial_lastname "LastName" \
                --port 61000
```

### Glossary/Reference

`<worker_path> = /home/cryosparc_user/software/cryosparc/cryosparc_worker`

* the **full** path to the worker directory, which was downloaded and extracted in step c)

  To get the full path, `cd` into the `cryosparc_worker` directory and run the command: `pwd -P`

`<ssd_path> = /scratch/cryosparc_cache`

* path on the worker node to a writable directory residing on the local SSD (For more information on SSD usage in cryoSPARC, see X)
* this is optional, and if omitted, specify the `--nossd` option to indicate that the worker node does not have an SSD

`<initial_email> = someone@structura.bio`

* login email address for first CryoSPARC webapp account
* this will become an admin account in the user interface

`<initial_username> = "FirstName LastName"`

* login username of the initial admin account to be created
* ensure the name is quoted

`<initial_firstname> = "FirstName"`

* given name of the initial admin account to be created
* ensure the name is quoted

`<initial_lastname> = "LastName"`

* surname of the initial admin account to be created
* ensure the name is quoted

`<initial_password> = Password123`

* temporary password that will be created for the `<user_email>` account
* Note that if this argument is not specified, a silent input prompt will be provided

`<port_number> = 61000`

* The base port number for this CryoSPARC instance. Do not install cryosparc master on the same machine multiple times with the same port number - this can cause database errors. **Choose the base port number carefully to**
  * avoid conflicts with other applications
  * avoid conflicts with other CryoSPARC instances that may be running on the same computer

{% hint style="info" %}
Versions of CryoSPARC prior to v4.4.0 also require

`--cudapath <cuda path>`

where \<cuda path> corresponds to the CUDA installation directory, such as

`/opt/cuda-11.8,` that *contains* the `bin/` and `lib64/` subdirectories on the worker node, *not* the `cuda-11.8/bin/` subdirectory.
{% endhint %}
{% endtab %}

{% tab title="Master Node Only" %}

### Master Node CryoSPARC Installation

```bash
cd cryosparc_master

./install.sh --license $LICENSE_ID \
             --hostname <master_hostname> \
             --dbpath <db_path> \
             --port <port_number> \
             [--insecure] \
             [--allowroot] \
             [--yes]
```

**Example command execution**

```bash
./install.sh --license $LICENSE_ID \
             --hostname cryoem.cryosparcserver.edu \
             --dbpath /u/cryosparcuser/cryosparc/cryosparc_database \
             --port 61000
```

#### Start CryoSPARC

```
./bin/cryosparcm start
```

#### Create the first user

```
cryosparcm createuser --email "<user email>" \
                      --password "<user password>" \
                      --username "<login username>" \
                      --firstname "<given name>" \
                      --lastname "<surname>"
```

### Glossary/Reference

`--license $LICENSE_ID`

* the LICENCE\_ID variable exported in step a)

`<master_hostname> = your.cryosparc.hostname.com`

* the hostname of the server where the master is to be installed on

`<port_number> = 61000`

* The base port number for this CryoSPARC instance. Do not install cryosparc master on the same machine multiple times with the same port number - this can cause database errors. **Choose the base port number carefully to**
  * avoid conflicts with other applications
  * avoid conflicts with other CryoSPARC instances that may be running on the same node

`<db_path> = /u/cryosparcuser/cryosparc/cryosparc_database`

* the absolute path to a folder where the CryoSPARC database is to be installed. Ensure this location is in a readable and writeable location. The folder will be created if it doesn't already exist

`<initial_email> = someone@structura.bio`

* login email address for first CryoSPARC webapp account
* this will become an admin account in the user interface

`<initial_username> = "FirstName LastName"`

* login username of the initial admin account to be created
* ensure the name is quoted

`<initial_firstname> = "FirstName"`

* given name of the initial admin account to be created
* ensure the name is quoted

`<initial_lastname> = "LastName"`

* surname of the initial admin account to be created
* ensure the name is quoted

`<user_password> = Password123`

* temporary password that will be created for the `<user_email>` account

`--insecure`

* *\[optional]* specify this option to ignore SSL certificate errors when connecting to HTTPS endpoints. This is useful if you are behind an enterprise network using SSL injection.

`--allowroot`

* *\[optional]* if you run the installation function with root privileges, it will fail. Specify this option to force CryoSPARC to be installed as the root user.

`--yes`

* *\[optional]* do not ask for any user input confirmations
  {% endtab %}
  {% endtabs %}

### \[Optional] Re-load your `bashrc`**.** This will allow you to run the `cryosparcm` management script from anywhere in the system:

```
source ~/.bashrc
```

### **Access the User Interface**

After completing the above, navigate your browser to `http://<workstation_hostname>:<base_port_number>` to access the CryoSPARC user interface.

{% content-ref url="/pages/-M8D1uoTdonr7yVA223d" %}
[Accessing the CryoSPARC User Interface](/setup-configuration-and-management/how-to-download-install-and-configure/accessing-cryosparc)
{% endcontent-ref %}

{% hint style="warning" %}
If you were following the **Single Workstation** install above, your installation is now complete.
{% endhint %}

## Install The **`cryosparc_worker`** Package

Log onto the worker node (or a cluster worker node if installing on a cluster) as the `cryosparcuser`, and run the installation command.

{% hint style="info" %}
Installing `cryosparc_worker` requires a Nvidia GPU and Nvidia Driver version 520.61.05 or newer
{% endhint %}

### GPU Worker Node CryoSPARC Installation

```
cd cryosparc_worker

./install.sh --license $LICENSE_ID [--yes]
```

### Worker Installation Glossary/Reference

`--license $LICENSE_ID`

* The LICENCE\_ID exported in the "Export your License ID as an environment variable" step

`--yes`

* *\[optional]* do not ask for any user input confirmations

{% hint style="info" %}
Versions of CryoSPARC prior to v4.4.0 also require

`--cudapath <cuda path>`

such as

```
./install.sh --license $LICENSE_ID --cudapath /opt/cuda-11.8
```

* Path to the CUDA installation directory on the worker node
* **Note**: this path should not be the `cuda/bin/` sub directory, but the CUDA directory that contains both `bin/` and `lib64/` subdirectories
  {% endhint %}

## Connecting A Worker Node

### Connect A Managed Worker to CryoSPARC

Ensure cryoSPARC is running when connecting a worker node. Log into the worker node and run the connection function (substitute the placeholders with your own configuration parameters; details below):

```bash
cd cryosparc_worker

./bin/cryosparcw connect --worker <worker_hostname> \
                         --master <master_hostname> \
                         --port <port_num> \
                         --ssdpath <ssd_path> \
                         [--update] \
                         [--sshstr <custom_ssh_string> ] \
                         [--nogpu] \
                         [--gpus <0,1,2,3> ] \
                         [--nossd] \
                         [--ssdquota <ssd_quota_mb> ] \
                         [--ssdreserve <ssd_reserve_mb> ] \
                         [--lane <lane_name> ] \
                         [--newlane]
```

#### Worker Connection Glossary/Reference

<pre class="language-bash"><code class="lang-bash">--worker &#x3C;worker_hostname> = worker.cryosparc.hostname.com
# [required] hostname of the worker node to connect to your 
# cryoSPARC instance

--master &#x3C;master_hostname> = your.cryosparc.hostname.com
# [required] hostname of the server where the master is 
# installed on

--port &#x3C;port_number> = 61000
# [required] base port number for the cryoSPARC instance you 
# are connecting the worker node to

--ssdpath &#x3C;ssd_path>
# [optional] path to directory on local SSD
# replace this with --nossd to connect a worker node without an SSD

--update
# [optional] use when updating and existing configuration

--sshstr &#x3C;custom_ssh_string>
# [optional] custom SSH connection string such as user@hostname 

--cpus 4
# [optional] enable this number of CPU cores 

<strong>--nogpu
</strong># [optional] connect worker with no GPUs 

--gpus 0,1,2,3
# [optional] enable specific GPU devices only 

# For advanced configuration, run the gpulist command:
# $ bin/cryosparcw gpulist

# Detected 4 CUDA devices.

#   id           pci-bus  name
#   ---------------------------------------------------------------
#       0      0000:42:00.0  Quadro GV100
#       1      0000:43:00.0  Quadro GV100
#       2      0000:0B:00.0  Quadro RTX 5000
#       3      0000:0A:00.0  GeForce GTX 1080 Ti
#   ---------------------------------------------------------------
# This will list the available GPUs on the worker node, and their 
# corresponding numbers. Use this list to decide which GPUs you wish 
# to enable using the --gpus flag, or leave this flag out to enable all GPUs.

--nossd
# [optional] connect a worker node with no SSD 

--ssdquota &#x3C;ssd_quota_mb>
# [optional] quota of how much SSD space to use (MB) 

--ssdreserve &#x3C;ssd_reserve_mb>
# [optional] minimum free space to leave on SSD (MB) 

--lane &#x3C;lane_name>
# [optional] name of lane to add worker to 

--newlane
# [optional] force creation of a new lane if the lane specified by --lane 
# does not exist
</code></pre>

### Update a Managed Worker Configuration

To update an existing managed worker configuration, use `cryosparcw connect` with the `--update` flag and the field you would like to update.

For example:

```bash
cd cryosparc_worker
./bin/cryosparcw connect --worker <worker_hostname> \
                         --master <master_hostname> \
                         --port <port_num> \
                         --update
                         --ssdquota 500000
```

## Connect a Cluster to CryoSPARC

CryoSPARC is designed to interface with HPC cluster systems so that jobs can be launched on shared resources that may be used for multiple purposes and by multiple users. CryoSPARC supports most cluster schedulers, including SLURM, SGE and PBS. CryoSPARC architecture when connecting to a cluster is described [here](https://guide.cryosparc.com/setup-configuration-and-management/hardware-and-system-requirements#clusters). We provide examples of [cluster integration scripts](https://guide.cryosparc.com/setup-configuration-and-management/how-to-download-install-and-configure/cryosparc-cluster-integration-script-examples) which can be adapted as required for the specific needs of a particular cluster system.It is also possible to use [custom values in cluster submission scripts](/setup-configuration-and-management/software-system-guides/guide-configuring-custom-variables-for-cluster-job-submission-scripts), before a job is queued.

### Connecting CryoSPARC to a cluster during installation

Once the `cryosparc_worker` package is installed, the cluster must be registered in the CryoSPARC database. Registration is performed by first creating two files, `cluster_info.json` and `cluster_script.sh` in the current working directory, populating those files, and then running the [command](/setup-configuration-and-management/management-and-monitoring-v5.0/cryosparcm-reference-v5.0#cryosparcm-cluster):

```
cryosparcm cluster connect
```

* the cluster information file `cluster_info.json` contains template strings used to construct cluster commands (e.g., qsub, qstat, qdel etc., or their equivalents for your system)
* the cluster script template file `cluster_script.sh` contains a [jinja2](https://jinja.palletsprojects.com/en/2.10.x/) template string corresponding to a cluster job script.

Examples of the contents of these two files can be found here:

{% content-ref url="/pages/-M7DHIJv291IAXg6Tkl\_" %}
[CryoSPARC Cluster Integration Script Examples](/setup-configuration-and-management/how-to-download-install-and-configure/cryosparc-cluster-integration-script-examples)
{% endcontent-ref %}

{% hint style="info" %}
Creation and/or modification of the `cluster_info.json` and `cluster_script.sh` files is only one step in the configuration of a cluster lane for CryoSPARC jobs. To become effective, information in these files needs to be [registered in the CryoSPARC database](#register-the-cluster-lane-in-the-cryosparc-database) using the `cryosparcm cluster connect` command.
{% endhint %}

For the submission and managment of CryoSPARC cluster jobs, cluster information and script templates are read from the database, *not* from the `cluster_info.json` and `cluster_script.sh` files. The `jinja2` template engine renders the actual cluster management commands as well as the submission scripts for each CryoSPARC cluster job.

### Variables available in `cluster_info.json`

`name`: string, **required**

* Unique name for the cluster to be connected (multiple clusters can be connected).
* This will be the name of the lane users will see when queuing jobs in CryoSPARC.![](/files/mYXwyKSbvXWLkoeElLoy)

`worker_bin_path`: string, **required**

* Absolute path on cluster nodes to the `cryosparcw` script.

`cache_path`: string, optional

* Absolute path on cluster nodes that is a writable location on the local SSD on each cluster node. This might be `/scratch` or similar.
* This path **must** be the same on all cluster nodes. See [Guide: Installation Testing with cryosparcm test](/setup-configuration-and-management/software-system-guides/guide-installation-testing-with-cryosparcm-test#cryosparcm-test-workers) for instructions on how to verify that CryoSPARC is able to successfully write to the cache path.
* If you plan to use the cluster nodes without an SSD, you can omit this field.

`cache_reserve_mb`:, integer, optional

* The size (in MB) to initially reserve for the cache on the SSD. This value is 10GB by default, which means CryoSPARC will always leave at least 10GB of free space on the SSD.

`cache_quota_mb`: integer, optional

* The maximum size (in MB) to use for the cache on the SSD.

`send_cmd_tpl`: string, **required**

* Used to send a cluster management command to be executed by a cluster node (in case the cryosparc master is not able to directly use cluster management commands).
* If your cryosparc master node is able to directly use cluster management commands (e.g., `qsub`, etc.,) then this string can be just `"{{ command }}"`.

`qsub_cmd_tpl`: string, **required**

* The cluster management command used to submit a job to the cluster, where the job is defined in the cluster script located at `{{ script_path_abs }}`.
* This string can also use any of the variables defined in cluster\_script.sh that are available when the job is scheduled (e.g., `cryosparc_username`, `project_uid`, etc.,).

`qstat_cmd_tpl`: string, **required**

* The cluster management command that will report back the status of cluster job with its ID `{{ cluster_job_id }}`.

`qdel_cmd_tpl`: string, **required**

* The cluster management command that will kill and remove the cluster job (using `{{ cluster_job_id }}`) from the queue.

`qinfo_cmd_tpl`: string, **required**

* The cluster management command to retrieve general cluster information.

Along with the configuration variables above, a complete cluster configuration requires a template cluster submission script that needs to be prepared in a file named `cluster_script.sh`. The script must send jobs into your cluster scheduler queue and mark them with the appropriate hardware requirements. The CryoSPARC internal scheduler submits jobs with this script as their inputs become ready. The following variables are available for use used within a cluster submission script template. When starting out, example templates may be generated with the commands explained below.

### The Basic Set of Variables Available in `cluster_script.sh` <a href="#cluster-template-basic-variables" id="cluster-template-basic-variables"></a>

```
{{ script_path_abs }}    # absolute path to the generated submission script
{{ run_cmd }}            # complete command-line string to run the job
{{ num_cpu }}            # number of CPUs needed
{{ num_gpu }}            # number of GPUs needed.
{{ ram_gb }}             # amount of RAM needed in GB
{{ job_dir_abs }}        # absolute path to the job directory
{{ project_dir_abs }}    # absolute path to the project dir
{{ job_log_path_abs }}   # absolute path to the log file for the job
{{ worker_bin_path }}    # absolute path to the cryosparc worker command
{{ run_args }}           # arguments to be passed to cryosparcw run
{{ project_uid }}        # uid of the project
{{ job_uid }}            # uid of the job
{{ job_creator }}        # name of the user that created the job (may contain spaces)
{{ cryosparc_username }} # cryosparc username of the user that created the job (usually an email)
{{ job_type }}           # CryoSPARC job type
```

{% hint style="info" %}
The CryoSPARC scheduler does not control GPU allocation when spawning jobs on a cluster. The number of GPUs required is provided as a template variable. Either your submission script or your cluster scheduler is responsible for assigning GPU device indices to each job spawned based on the provided variable. The CryoSPARC worker processes that use one or more GPUs on a cluster simply use device 0, then 1, then 2, etc.
{% endhint %}

The cluster script template can be further customized with additional variables, as described in

{% content-ref url="/pages/5zikU1RIs7uHyIpgV96G" %}
[Guide: Configuring Custom Variables for Cluster Job Submission Scripts](/setup-configuration-and-management/software-system-guides/guide-configuring-custom-variables-for-cluster-job-submission-scripts)
{% endcontent-ref %}

### Register the Cluster Lane in the CryoSPARC Database

The `cryosparcm cluster connect` command activates a new or modified cluster lane configuration by uploading information from the files `cluster_info.json` and `cluster_script.sh` to the CryoSPARC database.

{% hint style="info" %}
The`cryosparcm cluster connect` command attempts reading `cluster_info.json` and `cluster_script.sh` from the current working directory.
{% endhint %}

{% hint style="warning" %}
The `cryosparcm cluster connect` command will overwrite an existing database record if the existing record's `name` matches the `"name"` value inside `cluster_info.json`.
{% endhint %}

```bash
cryosparcm cluster example <cluster_type>
# dumps out config and script template files to current working directory
# examples are available for pbs and slurm schedulers but others should 
# be very similar

cryosparcm cluster dump <name>
# dumps out existing config and script to current working directory

cryosparcm cluster connect
# connects new or updates existing cluster configuration, 
# reading cluster_info.json and cluster_script.sh from the current directory, 
# using the name from cluster_info.json

cryosparcm cluster remove <name>
# removes a cluster configuration from the scheduler
```

### Test the Registered Cluster Lane

The success of the cluster lane's registration may be tested by queuing a CryoSPARC job to the cluster. For example, one may for this purpose:

* clone an existing job that was previously completed while the CryoSPARC instance ran on the current CryoSPARC version and queue the clone
* create an [Extensive Validation](/setup-configuration-and-management/software-system-guides/tutorial-verify-cryosparc-installation-with-the-extensive-workflow-sysadmin-guide) job and queue it

to the newly connected cluster lane

### Update a Cluster Lane Configuration

To update an existing cluster integration, call the `cryosparcm cluster connect` command with the updated `cluster_info.json` and `cluster_script.sh` in the current working directory.

{% hint style="info" %}
Note that the `name` field from `cluster_info.json` must be the same in the cluster configuration to update
{% endhint %}

If you don't already have the `cluster_info.json` and `cluster_script.sh` files in your current working directory, you can get them by running the command `cryosparcm cluster dump <name>`

### Add Additional Cluster Lanes

Additional lanes for a cluster can be added to increase options for submission parameters when submitting jobs to the cluster. The following steps can be followed to create a new lane from an existing cluster configuration:

Prerequisites:

* a folder containing the cluster setup files - `cluster_folder`
* cluster configuration files - `cluster_folder/cluster_info.json` and `cluster_folder/cluster_script.sh`

Steps:

1. create a new directory `cluster_folder_new_lane` parallel to `cluster_folder`
2. copy the cluster configuration files `cluster_folder/cluster_info.json` and `cluster_folder/cluster_script.sh` to `cluster_folder_new_lane`
3. change the `name` field in `cluster_folder_new_lane/cluster_info.json` to the desired name of the new lane
4. edit `cluster_folder_new_lane/cluster_info.json` and `cluster_folder_new_lane/cluster_script.sh` with any lane configuration changes
5. set `cluster_folder_new_lane` as the current working directory
6. connect the new cluster lane with `cryosparcm cluster connect` (this must be run with `cluster_folder_new_lane` as the working directory)

Note: The folder names `cluster_folder` and `cluster_folder_new_lane` can be changed to more descriptive names if desired, such as the name of the lane.


# CryoSPARC Cluster Integration Script Examples

Examples of cluster\_info.json and cluster\_script.sh scripts for various cluster workload managers

CryoSPARC can integrate with cluster scheduler systems. This page contains examples of integration setups.

Due to the many variations of cluster scheduler systems and their configurations, the examples here will need to be modified for your own specific use case.

For information on what each variable means, see [Downloading and Installing CryoSPARC](/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc#connect-a-cluster-to-cryosparc).

## GPU Resource Management

When CryoSPARC launches a job to the cluster, the number of GPUs requested by the user is used in the submission script, but CryoSPARC does not know which specific devices it will be allocated and therefore each job simply tries to use GPUs with device numbers starting at zero. For example, a 2-GPU job will try to use GPUs \[0, 1] when submitted to a cluster. It is the responsibility of the cluster system to correctly allocate requested GPU resources to CryoSPARC jobs while insulating those allocated resources from interference by other jobs. For this purpose, the SLURM scheduler, for example, can combine [Generic Resource (GRES)](https://slurm.schedmd.com/gres.html) management with (Linux) [cgroup](https://slurm.schedmd.com/cgroups.html)[ controls](https://slurm.schedmd.com/cgroups.html).

## A SLURM Example

{% code title="cluster\_info.json" %}

```json
{
    "name": "slurm-lane1",
    "worker_bin_path": "/path/to/cryosparc_worker/bin/cryosparcw",
    "send_cmd_tpl": "{{ command }}",
    "qsub_cmd_tpl": "/opt/slurm/bin/sbatch {{ script_path_abs }}",
    "qstat_cmd_tpl": "/opt/slurm/bin/squeue -j {{ cluster_job_id }}",
    "qdel_cmd_tpl": "/opt/slurm/bin/scancel {{ cluster_job_id }}",
    "qinfo_cmd_tpl": "/opt/slurm/bin/sinfo"
}
```

{% endcode %}

{% code title="cluster\_script.sh" %}

```django
#!/usr/bin/env bash

#SBATCH --job-name cryosparc_{{ project_uid }}_{{ job_uid }}
#SBATCH --cpus-per-task={{ num_cpu }}
#SBATCH --gres=gpu:{{ num_gpu }}
#SBATCH --mem={{ ram_gb|int }}G
#SBATCH --comment="created by {{ cryosparc_username }}"
#SBATCH --output={{ job_dir_abs }}/{{ project_uid }}_{{ job_uid }}_slurm.out
#SBATCH --error={{ job_dir_abs }}/{{ project_uid }}_{{ job_uid }}_slurm.err

{{ run_cmd }}
```

{% endcode %}

## A PBS Example

{% code title="cluster\_info.json" %}

```json
{
    "name" : "pbscluster",
    "worker_bin_path" : "/path/to/cryosparc_worker/bin/cryosparcw",
    "cache_path" : "/path/to/local/SSD/on/cluster/nodes",
    "send_cmd_tpl" : "ssh loginnode {{ command }}",
    "qsub_cmd_tpl" : "qsub {{ script_path_abs }}",
    "qstat_cmd_tpl" : "qstat -as {{ cluster_job_id }}",
    "qdel_cmd_tpl" : "qdel {{ cluster_job_id }}",
    "qinfo_cmd_tpl" : "qstat -q"
}
```

{% endcode %}

{% code title="cluster\_script.sh" %}

```django
#!/bin/bash

#PBS -N cryosparc_{{ project_uid }}_{{ job_uid }}
#PBS -l select=1:ncpus={{ num_cpu }}:ngpus={{ num_gpu }}:mem={{ (ram_gb*1000)|int }}mb:gputype=P100
#PBS -o {{ job_dir_abs }}/cluster.out
#PBS -e {{ job_dir_abs }}/cluster.err

{{ run_cmd }}
```

{% endcode %}

## A Gridengine Example

{% code title="cluster\_info.json" %}

```json
{
    "name" : "ugecluster",
    "worker_bin_path" : "/u/cryosparcuser/cryosparc/cryosparc_worker/bin/cryosparcw",
    "cache_path" : "/scratch/cryosparc_cache",
    "send_cmd_tpl" : "{{ command }}",
    "qsub_cmd_tpl" : "qsub {{ script_path_abs }}",
    "qstat_cmd_tpl" : "qstat -j {{ cluster_job_id }}",
    "qdel_cmd_tpl" : "qdel {{ cluster_job_id }}",
    "qinfo_cmd_tpl" : "qstat -q default.q"
}
```

{% endcode %}

{% code title="cluster\_script.sh" %}

```django
#!/bin/bash

## What follows is a simple UGE script:
## Job Name
#$ -N cryosparc_{{ project_uid }}_{{ job_uid }}

## Number of CPUs (select 1 CPU always, and oversubscribe as GPU is per core value)
##$ -pe smp {{ num_cpu }}
#$ -pe smp 1

## Memory per CPU core
#$ -l m_mem_free={{ (ram_gb)|int }}G

## Number of GPUs 
#$ -l gpu_card={{ num_gpu }}

## Time limit 4 days
#$ -l h_rt=345600

## STDOUT/STDERR
#$ -o {{ job_dir_abs }}/cluster.out
#$ -e {{ job_dir_abs }}/cluster.err
#$ -j y

## Number of threads
export OMP_NUM_THREADS={{ num_cpu }}

echo "HOSTNAME: $HOSTNAME"

{{ run_cmd }}
```

{% endcode %}


# Accessing the CryoSPARC User Interface

Viewing the user interface locally and from home

The CryoSPARC user interface is served by a web server running on the same computer where CryoSPARC is installed, at the base port specified during installation. This web server is responsible for displaying datasets, experiments, streaming real time results, user accounts, updating, etc.

When you install CryoSPARC, you will be shown details on how to access the interface using the configuration you've provided.

```
cryosparcuser@csserver:~$ cryosparc_master/bin/cryosparcm start
Starting CryoSPARC System master process...
CryoSPARC is not already running.
configuring database...
    configuration complete
database: started
database OK
command_core: started
    command_core connection succeeded
    command_core startup successful
command_vis: started
command_rtp: started
    command_rtp connection succeeded
    command_rtp startup successful
app: started
app_api: started
-----------------------------------------------------

CryoSPARC master started. 
 From this machine, access CryoSPARC and CryoSPARC Live at
    http://localhost:61000

 From other machines on the network, access CryoSPARC and CryoSPARC Live at
    http://csserver.lab:61000


Startup can take several minutes. Point your browser to the address
and refresh until you see the CryoSPARC web interface.
```

Given the example above, one may connect to the CryoSPARC user interface

* using a browser running on the CryoSPARC master computer: with URL `http://localhost:61000`
* using a browser running on the same network as the CryoSPARC master computer: with URL `http://csserver.lab:61000`
* using a browser running on another network, the URL depends on the network access method (see below)

![Typical network setup for a CryoSPARC user](/files/-M8D4cEhunEvSRgSOVgA)

When you are working from a remote network, you will usually not have direct access to the master node to use CryoSPARC as you usually would.

Often, the master CryoSPARC server may be behind a firewall, within a local network (LAN) at your institution. Only other machines that are on the same local network can connect to the master server at port 61000.

## VPN Access

Most institutions offer Virtual Private Network (VPN) capability which can allow you to connect to the institution's local network as if you are physically present at the office. There are different types of VPN connections, but most will allow you, once logged in, to connect to the CryoSPARC master server as you usually would, using your browser.

In some cases, your VPN may only allow certain types of connections, or your institution may allow for access over only some secure ports to your CryoSPARC master server, without a VPN log in. In both of these cases, if you are able to find a way to connect to your CryoSPARC master server using `SSH`, it is still possible to use CryoSPARC, even if you cannot connect to port 61000 as you usually would.

## SSH Access and Tunneling

When you want to access CryoSPARC from home or elsewhere to be able to run jobs and view results, it can be convenient to connect to the web server via an SSH tunnel. SSH tunneling is a method of transporting arbitrary networking data over an encrypted SSH connection.

SSH is a standard for secure remote logins and file transfers over untrusted networks. It also provides a way to secure the data traffic of any given application using port forwarding, basically tunneling any TCP/IP port over SSH. This means that the application data traffic is directed to flow inside an encrypted SSH connection so that it cannot be eavesdropped or intercepted while it is in transit. Source: [SSH Tunnel](https://www.ssh.com/ssh/tunneling)

![You may need to use a Virtual Private Network (VPN) client to connect to your institution's VPN in order to access the local network.](/files/-M8D4rCSLMPKRN2cc76q)

## SSH Local Port Forwarding

Supposing an example scenario where

* CryoSPARC is installed on a computer with hostname csserver.lab
* the [web application port](https://guide.cryosparc.com/setup-configuration-and-management/cryosparc-installation-prerequisites#id-4.-open-tcp-ports) is configured to "listen" at port number 61000
* you have ssh access to the `myname` account on csserver.lab
* the port 62222 on the "local" computer that runs your web browser is not currently in use

you can establish a tunnel between the *local* computer and the CryoSPARC server with the command

```
ssh -L 62222:localhost:61000 myname@csserver.lab
```

If you do not have ssh access to the CryoSPARC server, but do have ssh access to another computer (sshserver.lab for example) that itself can access the web application port on the CryoSPARC server, you can instead establish the tunnel with the command

```
ssh -L 62222:csserver.lab:61000 myname@sshserver.lab
```

In both examples, the first port number

* is a freely chosen port on the *local* ("browser") computer that must not already be in use
* determines the `<portnumber>` portion in the `http://localhost:<portnumber>` URL to which you will point your browser.

The second port number in both examples is prescribed by the web application port number configured during CryoSPARC installation. Based on the two ssh examples above, you would point your browser to `http://localhost:62222`

![CryoSPARC UI login page](/files/4VThpwWGLuQuMTtwkBlX)

## More complex SSH requirements and configurations

If you need to hop over one or more "jump" hosts to access the CryoSPARC server or alternative "tunnel" server, please refer the [OpenSSH wikibook](https://en.wikibooks.org/wiki/OpenSSH/Cookbook/Proxies_and_Jump_Hosts#Passing_Through_One_or_More_Gateways_Using_ProxyJump) for suggested `~/.ssh/config` configurations.

## Reverse Proxy

Refer to the following guide for more information on hosting the CryoSPARC web application via a reverse proxy server:

{% content-ref url="/pages/aPAVqd3gBmaY2wQNi34J" %}
[(Optional) Hosting CryoSPARC Through a Reverse Proxy](/setup-configuration-and-management/how-to-download-install-and-configure/optional-hosting-cryosparc-through-a-reverse-proxy)
{% endcontent-ref %}

## Appendix

### Appendix A: Setting up password-less SSH access to a remote workstation

Set up SSH keys for password-less access (only if you currently need to enter your password each time you ssh into the compute node).

1. If you do not already have SSH keys generated on your local machine, use `ssh-keygen` to do so. Open a terminal prompt on your local machine, and enter:

   ```
   ssh-keygen -t rsa -N "" -f $HOME/.ssh/id_rsa
   ```

   *Note: this will create an RSA key-pair with no passphrase.*<br>
2. Copy the RSA public key to the remote compute node for password-less login:

   ```
   ssh-copy-id remote_username@remote_hostname
   ```

   *Note: `remote_username` and `remote_hostname` are your username and the hostname that you use to SSH into your compute node. This step will ask for your password.*

### Appendix B: Using SSH Forwarding with compression to reduce data usage

Supply `-C` to the port tunnelling command to request compression of all data. This can help when downloading maps from the CryoSPARC UI, as masks can be greatly compressed. From `man ssh`:

```
-C           Requests compression of all data (including stdin, stdout, stderr, 
             and data for forwarded X11, TCP and UNIX-domain connections).  
             The compression algorithm is the same used by gzip(1), and the 
             “level” can be controlled by the CompressionLevel option for 
             protocol version 1.  Compression is desirable on modem lines and 
             other slow connections, but will only slow down things on fast 
             networks.
```

For example:

```
ssh -N -f -L localhost:62222:localhost:61000 remote_hostname -C
```

### Appendix C: Using Hardware Accelerated OpenSSH Ciphers

If your system supports Intel or AMD AES-NI, you can take advantage of hardware accelerated ciphers that dramatically improve the performance of your SSH connection. To find out if your system supports this, [follow this tutorial.](https://www.cyberciti.biz/faq/how-to-find-out-aes-ni-advanced-encryption-enabled-on-linux-system/)

If your system has these features enabled, supply the argument `-o Ciphers=aes128-gcm@openssh.com` or `-o Ciphers=aes256-gcm@openssh.com` (depending on what your system supports, but AES 256 is preferred) to the port forwarding command. For example:

```
ssh -N -f -L localhost:62222:localhost:61000 remote_hostname -C -o Ciphers=aes256-gcm@openssh.com
```

### Appendix D: Custom SSL Certificate Authority Bundle

cryoSPARC requires internet access from the main process to verify your license and perform updates. At minimum, CryoSPARC should have access to our license server at `https://get.cryosparc.com/`.

On some older systems, or if your system is behind a HTTP proxy, CryoSPARC may have trouble getting the required SSL certificates to validate this requires. If you have a Certificate Authority (CA) bundle on your system, you may specify its path for CryoSPARC to use and apply.

Add the following line to `cryosparc_master/config.sh` (substitute `/path/to/cabundle` with the path to the CA bundle on your system):

```bash
export REQUESTS_CA_BUNDLE="/path/to/cabundle"
```


# (Optional) Hosting CryoSPARC Through a Reverse Proxy

As discussed in [Accessing the CryoSPARC User Interface](/setup-configuration-and-management/how-to-download-install-and-configure/accessing-cryosparc), there are various ways in which users can access the CryoSPARC web interface such as through a VPN connection or SSH tunnel. If you would like to host the CryoSPARC interface in a secure manner at a predictable URL, this can be done through a reverse proxy server.

Reverse proxy servers allow for more control over how a user accesses a web application interface over other methods. By controlling incoming network traffic it is able to host the application at a static URL (for example `https://cryosparc.institution.edu`) and ensure all correspondence is secured via HTTPS.

The method in which you host CryoSPARC through a reverse proxy is similar to hosting any other web application. However, the following serve as our recommended minimum requirements:

* All incoming traffic should be served through HTTPS (via a SSL certificate)
* HTTPS traffic requires a valid SSL certificate provided by a certificate authority (CA) for the domain in which you are hosting the interface.
* If the server listens for incoming HTTP traffic, forward all connections to a more secure protocol (HTTPS)
* Ensure traffic is also mediated by an organization-level authentication barrier (for example single sign-on). CryoSPARC should *not* be served via the public internet without any additional authentication checks.

There are many ways to generate a SSL certificate for your domain, however, this will most likely be specific to your institution or organization. If you're unsure of how to generate a SSL certificate for your private network, please consult with your system or network administrator for guidance.

Each institution or private network can have a specific setup requiring custom rules and/or proxy configuration considerations. Generally the example configurations below should be compatible with common reverse proxy installations. Please consult with your system or network administrator for guidance regarding institution-specific protocols for reverse proxy hosting.

The following section will provide example configuration files for common reverse proxy servers given CryoSPARC is running on base port `61000`.

## NGINX

This [NGINX](https://nginx.org/en/) configuration takes advantage of [authenticated origin pulls](https://developers.cloudflare.com/ssl/origin-configuration/authenticated-origin-pull/explanation/) for an added layer of security between the reverse-proxy and a downstream proxy/load balancer.

```
server {
  listen 80;
  listen [::]:80;
  server_name private.domain.dev;
  return 302 https://$server_name$request_uri;
}

server {
  # SSL configuration
  listen 443 ssl http2;
  listen [::]:443 ssl http2;
  ssl        on;
  ssl_certificate         /etc/certs/domain.dev/origin-cert.pem;
  ssl_certificate_key     /etc/certs/domain.dev/private-key.pem;
  ssl_client_certificate  /etc/certs/domain.dev/origin-pull-ca.pem;
  ssl_verify_client on;

  server_name   private.domain.dev;
  access_log    /var/log/nginx/private.domain.dev.access.log;
  error_log     /var/log/nginx/private.domain.dev.error.log;

  location / {
    proxy_pass http://127.0.0.1:61000;
    proxy_http_version 1.1;
    proxy_set_header Upgrade $http_upgrade;
    proxy_set_header Connection 'upgrade';
    proxy_set_header X-Forwarded-For $remote_addr;

    proxy_request_buffering  off;
    proxy_buffering          off;
    client_max_body_size     0;
  }
}
```

An alternative configuration from our [Discussion Forum](https://discuss.cryosparc.com/t/how-to-run-cryosparc-on-an-https-url-using-an-ssl-certificate/2614/7) that serves the application over HTTPS and redirects incoming HTTP requests.

```
server {
  listen                80;
  server_name           <YOUR_URL>;

  access_log            /var/log/nginx/<YOUR_URL>.http.access.log;
  error_log             /var/log/nginx/<YOUR_URL>.http.error.log;

  location / {
    return 301 https://$server_name$request_uri;
  }
}


server {
  listen                443 ssl;

  server_name           <YOUR_URL>;

  access_log            /var/log/nginx/<YOUR_URL>.access.log;
  error_log             /var/log/nginx/<YOUR_URL>.error.log;

  location / {
    proxy_pass http://127.0.0.1:61000;
    proxy_http_version 1.1;
    proxy_set_header Upgrade $http_upgrade;
    proxy_set_header Connection 'upgrade';
    proxy_set_header X-Forwarded-For $remote_addr;

    proxy_request_buffering  off;
    proxy_buffering          off;
    client_max_body_size     0;
  }
  ssl_certificate /etc/letsencrypt/live/<YOUR_URL>/fullchain.pem; # managed by Certbot
  ssl_certificate_key /etc/letsencrypt/live/<YOUR_URL>/privkey.pem; # managed by Certbot
}
```

## Apache

The following is a simplified [Apache HTTP Server](https://httpd.apache.org/) configuration that illustrates the `RewriteRule`. For production use, we recommend HTTPS instead of HTTP. Additional configuration, not shown here, is required to enable HTTPS.

```
<VirtualHost *:80>
	ProxyRequests Off
	RewriteEngine on
	ProxyPass / http://localhost:61000/
	ProxyPassReverse / http://localhost:61000/
	RewriteCond %{HTTP:UPGRADE} ^WebSocket$ [NC]
	RewriteCond %{HTTP:CONNECTION} ^Upgrade$ [NC]
	RewriteRule .* ws://localhost:61000%{REQUEST_URI} [P]
</VirtualHost>
```


# Software Updates and Patches

How to get the latest CryoSPARC features and fixes or roll back to a previous version.

{% hint style="warning" %}
CryoSPARC v5 was released January 27, 2026. For updating from CryoSPARC v4 to CryoSPARC v5 for the first time, see the [CryoSPARC v5 compatibility requirements and update instructions here](/setup-configuration-and-management/software-system-guides/guide-updating-to-cryosparc-v5).
{% endhint %}

## Software updates

### New versions

When we release a new version of CryoSPARC, the Dashboard will display an update notification.

In CryoSPARC v4 and v5, the update notification is displayed at the top of the Dashboard:

<figure><img src="/files/FSQJrdV4QnlAhT2Pdaz3" alt=""><figcaption></figcaption></figure>

In CryoSPARC v3, the update notification is displayed in the footer:

![Look for the new version notification in the bottom status bar of the cryoSPARC dashboard.](/files/-M7DHJj3N_-VSzx2d2km)

You can [**sign up for the newsletter**](https://cryosparc.com/#newsletter) to receive an email when we release a new update or patch.

### Patches

We also periodically release patch updates for CryoSPARC, to address issues outside of normal updates. To learn more about patches, see: <https://guide.cryosparc.com/setup-configuration-and-management/software-updates#apply-patches>

{% embed url="<https://guide.cryosparc.com/setup-configuration-and-management/software-updates#apply-patches>" %}

## Before you update or downgrade

{% hint style="warning" %}
For updating from CryoSPARC v4 to CryoSPARC v5 for the first time, see the [CryoSPARC v5 compatibility requirements and update instructions here](/setup-configuration-and-management/software-system-guides/guide-updating-to-cryosparc-v5).
{% endhint %}

### 1. Confirm sufficient storage capacity

During the update, downloaded packages will require *additional* storage of approximately 6 gigabytes in the `cryosparc_master/` directory and 5 gigabytes in the `cryosparc_worker/` directory (for CryoSPARC version 4.7+).

If `cryosparc_master/` and/or `cryosparc_worker/` are stored on the same volume as the CryoSPARC database, one must additionally ensure that the downloaded update packages will not fill up storage needed for the database.

{% hint style="warning" %}
Failure to ensure sufficient storage may result in

* a failed software update
* a broken software installation
* when installation directories are stored on the same volume as the CryoSPARC database, a corrupt database
  {% endhint %}

### 2. Complete or \`kill\` running jobs

Before you update your instance, wait for all presently running jobs to complete, or kill them from the **Resource Manager**. We also **highly** recommend making a backup of your database as described below.

{% hint style="danger" %}
**IMPORTANT.** Before you update your instance, we recommend making a [**backup of your database.**](https://guide.cryosparc.com/setup-configuration-and-management/management-and-monitoring/cryosparcm#cryosparcm-backup)
{% endhint %}

In CryoSPARC v4.0+, you can use the Maintenance Mode feature to pause new jobs during an update:

{% content-ref url="/pages/ExdUPe4xFh3gYepmHpCq" %}
[Guide: Maintenance Mode and Configurable User Facing Messages](/setup-configuration-and-management/software-system-guides/guide-maintenance-mode-and-configurable-user-facing-messages)
{% endcontent-ref %}

### 3. Create a backup of the database

{% hint style="danger" %}
**IMPORTANT.** Before you update your instance, we recommend making a [**backup of your database.**](https://guide.cryosparc.com/setup-configuration-and-management/management-and-monitoring/cryosparcm#cryosparcm-backup)
{% endhint %}

### 4. *Completely* shutdown CryoSPARC (with confirmation)

{% hint style="warning" %}
Incomplete shutdown of the CryoSPARC instance during updates is a known cause for update failures. Follow the sequence for a [*complete* shutdown](/setup-configuration-and-management/troubleshooting#incomplete-cryosparc-shutdown) of the CryoSPARC instance.
{% endhint %}

## Checking for updates

{% hint style="warning" %}
Unless otherwise noted:

* Log in to the workstation or remote node where `cryosparc_master` is installed.
* Use the same non-root UNIX user account that runs the cryoSPARC process and was used to install cryoSPARC.
* Run all commands on this page in a terminal running `bash`
  {% endhint %}

Run this command to check for CryoSPARC updates.

```
cryosparcm update --check
```

This checks online for available updates and indicates whether an update is available.

```bash
$ cryosparcm update --check

CryoSPARC current version v2.15.0
          update starting on Wed Mar 18 12:09:52 EDT 2020

  current version v2.15.0
      new version v3.0.0

Update available!
```

You can also use `cryosparcm update --list` to get a full list of available versions (including old versions in case you would like to downgrade).

```
$ cryosparcm update --list

CryoSPARC current version v4.1.0
          update starting on Wed Mar 18 12:11:42 EDT 2020

Available versions:

v2.0.18
v2.0.20
v2.0.23
...
v4.0.2
v4.0.3
v4.1.0

To install a specific version, use 
    $ cryosparcm update --version=vXX.YY.ZZ[-branchname]
```

## Installing automatic updates

{% hint style="warning" %}
For updating from CryoSPARC v4 to CryoSPARC v5 for the first time, see the [CryoSPARC v5 compatibility requirements and update instructions here](/setup-configuration-and-management/software-system-guides/guide-updating-to-cryosparc-v5).
{% endhint %}

Perform the following actions when installing the latest version of CryoSPARC.

To begin automatic master and *non-cluster* worker updates with the newest available version of CryoSPARC, run

```bash
cryosparcm update
```

{% hint style="warning" %}
Cluster workers are *not* updated automatically. See the "Manual Cluster Updates" section below
{% endhint %}

This commands executes the following:

#### 1. Runs an automatic master update

* Downloads the new master (`cryosparc_master.tar.gz`) and worker (`cryosparc_worker.tar.gz`) update packages if `--skip-download` was not specified
* Shuts down the running CryoSPARC instance
* Extracts and installs the master release
* If dependencies have changed, automatically re-installs these

CryoSPARC releases include many compressed files; the extraction step may take several minutes on slower disks.

#### 2. Runs automatic worker updates

Once the master update is complete, master starts up and automatically updates registered workers:

* Transfers the worker release `cryosparc_worker.tar.gz` to each worker node via `scp`
* Extracts and installs the worker release
* Updates dependencies

If multiple standalone worker nodes are registered that all share the same worker installation, the update is only applied once.

### *Manual* Cluster Updates

Cluster installations **do not update automatically** because not all clusters have internet access on worker nodes.

Once the automatic update above is complete, navigate to the CryoSPARC master installation directory via command-line. Look for the latest downloaded worker release, named `cryosparc_worker.tar.gz` or`cryosparc2_worker.tar.gz`

Copy this file (via `scp`) to cluster worker's installation directory. It should be in the same directory as the `bin` and `deps` folders. Navigate to the installation directory and run

{% hint style="info" %}
If you're updating from cryoSPARC v2 to v3, the downloaded file is called`cryosparc2_worker.tar.gz`. Do not change this file name when you copy it into the worker directory.
{% endhint %}

```
bin/cryosparcw update
```

This updates the worker at the current location with the given release file.

{% hint style="info" %}
cryoSPARC does not allow running mismatched versions of master and worker releases. If you see this error:

```
Version mismatch! Worker and master versions are not the same. Please update.
```

Then re-install your cryoSPARC master and worker and check they are on the same version.
{% endhint %}

## Update or roll back/downgrade to a specific version

Follow this section to install or update/downgrade to a specific release of CryoSPARC.

Please see [**Before you update or downgrade**](#before-you-update-or-downgrade-complete-or-kill-running-jobs).

Steps are as described above, but with this command instead

```
cryosparcm update --version=vX.Y.Z
```

Use `cryosparcm update --list` to see the list of available versions. Substitute the `vX.Y.Z` in the command above with one of the results.

{% hint style="danger" %}
For instructions on updating an existing CryoSPARC v3.x instance to v4.0, please see: [Guide: Updating to CryoSPARC v4](/setup-configuration-and-management/software-system-guides/guide-updating-to-cryosparc-v4)
{% endhint %}

{% hint style="danger" %}
If your CryoSPARC instance is running v4.0.0 and later, the oldest version of CryoSPARC you can downgrade to is v3.4.0. For more information, see [Guide: Updating to CryoSPARC v4](/setup-configuration-and-management/software-system-guides/guide-updating-to-cryosparc-v4#downgrading)
{% endhint %}

## Forced update

Follow this section when a cryoSPARC install or update process fails part-way, or if CryoSPARC cannot start after following the [**Troubleshooting**](/setup-configuration-and-management/troubleshooting) steps.

This removes CryoSPARC and installs the latest available version, bypassing all file and dependency checks.

On the master node run

```
cryosparcm update --override
```

Then on each worker node run

```
bin/cryosparcw update --override
```

{% hint style="info" %}
You cannot specify a version to install when overriding the update manager. Only the latest version of cryoSPARC will be installed.
{% endhint %}

## Skip Downloading Update Packages During Update

Use the command line flag `--skip-download` with the `cryosparcm update` command to perform a full update of CryoSPARC using the update packages already inside the `cryosparc_master` directory. This command can be run after manually downloading the update packages using `cryosparcm update --download-only`.

For example:

`cryosparcm update --skip-download`

## Manually Download Update Packages

Use the command line flag `--download-only` with the `cryosparcm update` command to only download the update packages and not perform a full update of CryoSPARC. This can be helpful if your internet connection is slow, and you'd like to limit CryoSPARC's downtime. This option can be used to download any version of CryoSPARC available (use `cryosparcm update --list` to see the list of available versions).

You can then use `cryosparcm update --skip-download` to perform a full update using the downloaded update packages.

For example:

`cryosparcm update --download-only` # downloads the latest update packages

`cryosparcm update --version=v4.1.0 --download-only` # downloads the update packages corresponding to `v4.1.0`

## Verify successful update or installation (optional)

CryoSPARC provides two methods of verifying that all components of an installation are correctly working and set up.

The first method is to run `cryosparcm test install` and `cryosparcm test workers` via the command line. For more information, see [Guide: Installation Testing with cryosparcm test](/setup-configuration-and-management/software-system-guides/guide-installation-testing-with-cryosparcm-test)

The second method is to run the Extensive Workflow job. [See the Extensive Workflow guide here.](/setup-configuration-and-management/software-system-guides/tutorial-verify-cryosparc-installation-with-the-extensive-workflow-sysadmin-guide)

This automatic workflow executes all steps in the [T20S Introductory Tutorial](/guides-for-v3/cryo-em-data-processing-in-cryosparc-introductory-tutorial) and verifies the following system components:

* CryoSPARC system and license installation
* Worker/Cluster configuration
* GPU and CUDA driver installation
* SSD caching

## Apply Patches

We periodically releases patches for specific versions of CryoSPARC to fix bugs which do not require a full formal software update. [**Subscribe to the CryoSPARC Newsletter**](https://cryosparc.com/#newsletter) to receive an email when we release a patch.

{% hint style="info" %}
Patches for a specific CryoSPARC version are cumulative:

* One may apply the latest patch without having applied earlier patches.
* After application of the latest patch, application of an earlier patch is not needed and should not be attempted.

One may apply the latest patch for a given CryoSPARC version even if an earlier patch for that version has previously been applied.
{% endhint %}

To check for available patches, run

```
cryosparcm patch --check
```

Before applying patches, ensure CryoSPARC is running:

```
cryosparcm start
```

Apply the patch with one of the following strategies (table):

{% hint style="info" %}
If a cluster is connected, use the cluster installation instructions even if CryoSPARC was initially installed in workstation or master/worker mode
{% endhint %}

{% tabs %}
{% tab title="Single Workstation" %}
Automatically install all patches:

```
cryosparcm patch
```

{% endtab %}

{% tab title="Master Node" %}
Automatically install patches on the master and connected dedicated worker nodes

```
cryosparcm patch
```

{% endtab %}

{% tab title="Cluster" %}
From the master node, run

```
cryosparcm patch --download
```

This downloads master and worker tarballs to the `cryosparc_master` installation directory. Follow the resulting set of instructions for installing both patch files.

The instructions will involve the following:

* Install the master patch file with `cryosparcm patch --install`
* Copy or upload the downloaded `cryosparc_worker_patch.tar.gz` patch into the `cryosparc_worker` directory
* Inside the `cryosparc_worker` directory, run `bin/cryosparcw patch`

Depending on your cluster setup, either install the patch once in the `cryosparc_worker` directory shared by all cluster nodes or repeatedly for each cluster node that hosts an independent `cryosparc_worker` directory.
{% endtab %}
{% endtabs %}

Finally, restart CryoSPARC:

```
cryosparcm restart
```


# Management and Monitoring (≤v4.7)

Instructions for accessing and working in the CryoSPARC command line.

{% hint style="warning" %}
This page refers to CryoSPARC ≤v4.7.

For v5.0+, please see [https://github.com/cryoem-uoft/guide-beta/blob/master/setup-configuration-and-management/management-and-monitoring-v5.0](https://github.com/cryoem-uoft/guide-beta/blob/master/setup-configuration-and-management/management-and-monitoring-v5.0 "mention")
{% endhint %}

## Environment variables

Specify additional environment variables in the configuration files to augment CryoSPARC's low-level behaviour.

{% content-ref url="/pages/P3sVaHfkUFPsou1J5FCe" %}
[Environment variables (≤v4.7)](/setup-configuration-and-management/management-and-monitoring-4.7/environment-variables-v4.7)
{% endcontent-ref %}

## cryosparcm, cryosparcm cli and cryosparcw references

Workstations or master nodes with a `cryosparc_master` installation have access to `cryosparcm`, CryoSPARC's built-in [command-line](https://en.wikipedia.org/wiki/Command-line_interface) utility for all administrative, management and advanced usage tasks.

{% content-ref url="/pages/-M7DHIJyF6g-u3KTukdT" %}
[cryosparcm reference (≤v4.7)](/setup-configuration-and-management/management-and-monitoring-4.7/cryosparcm-4.7)
{% endcontent-ref %}

The `cryosparcm cli` command provides an extensive API for programmatically controlling cryoSPARC from the command-line.

{% content-ref url="/pages/-M7DHIJzq3PnUov4PO7G" %}
[cryosparcm cli reference (≤v4.7)](/setup-configuration-and-management/management-and-monitoring-4.7/cli-4.7)
{% endcontent-ref %}

Workstations or worker nodes with a `cryosparc_worker` installation have access to `cryosparcw`, a utility similar to `cryosparcm` for managing worker installations.

{% content-ref url="/pages/-M7DHIK5Xfhg7tPgHfzc" %}
[cryosparcw reference (≤v4.7)](/setup-configuration-and-management/management-and-monitoring-4.7/cryosparcw-4.7)
{% endcontent-ref %}


# Environment variables (≤v4.7)

(Advanced) Specify additional environment variables in the configuration files to augment CryoSPARC's low-level behaviour.

{% hint style="warning" %}
This page refers to CryoSPARC ≤v4.7.

For v5.0+, please see [Environment Variables (v5.0+)](/setup-configuration-and-management/management-and-monitoring-v5.0/environment-variables-v5.0)
{% endhint %}

To set or change one of these environment settings, add a new line to one the `config.sh` files with the following format (substitute `VARIABLE` and `VALUE` as indicated in the next sections):

```bash
export VARIABLE="VALUE"
```

Or set the value based on a different environment variable provided by the system:

```bash
export VARIABLE="${OTHER_VARIABLE}"
```

Which `config.sh` file you use depends on which variable you need to change. The variables available for each file are described below.

## cryosparc\_master/config.sh

**Note:** Restart CryoSPARC with `cryosparcm restart` after changing this file.

| Variable                                       | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                   | Default Value               |
| ---------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------- |
| `CRYOSPARC_CLIENT_TIMEOUT`                     | How many seconds to wait for network requests between CryoSPARC master/master processes before timing out with an error                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       | `300`                       |
| `CRYOSPARC_CLUSTER_JOB_MONITOR_INTERVAL`       | The amount of time to wait (in seconds) in between status updates for cluster jobs.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           | `10`                        |
| `CRYOSPARC_CLUSTER_JOB_MONITOR_MAX_RETRIES`    | The maximum amount of retries for cluster job status updates.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                 | `1000000`                   |
| `CRYOSPARC_DB_CONNECTION_TIMEOUT_MS`           | The amount of time to wait (in milliseconds) for the database to respond when starting it.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    | `20000`                     |
| `CRYOSPARC_DISABLE_IMPORT_ON_MASTER`           | In master/worker or cluster modes, Import Jobs always run on the default master node. Set this to `true` to allow running them on any node                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    | `false`                     |
| `CRYOSPARC_FORCE_HOSTNAME`                     | In master/worker or cluster modes, `cryosparcm` commands always run on the machine where `cryosparc_master` was initially installed. Set this to `true` to allow running on any machine                                                                                                                                                                                                                                                                                                                                                                                                                                                       | `false`                     |
| `CRYOSPARC_FORCE_USER`                         | `cryosparcm` commands must be run by the same UNIX user account that owns the `cryosparc_master` installation. Set this to `true` to allow running `cryosparcm` from any user account                                                                                                                                                                                                                                                                                                                                                                                                                                                         | `false`                     |
| `CRYOSPARC_HOSTNAME_CHECK`                     | Override the installed master hostname when checking that `cryosparcm` is running on the correct host. See also `CRYOSPARC_FORCE_HOSTNAME`                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    | -                           |
| `CRYOSPARC_IGNORE_HIDDEN_FILES`                | When checking that job directory is empty before running a job, set to `true` to ignore any filenames that being with `.`. Use this if your file system or OS automatically creates hidden files                                                                                                                                                                                                                                                                                                                                                                                                                                              | `false`                     |
| `CRYOSPARC_IO_URING`                           | Set to `false` to force-disable io\_uring for fast disk read operations during jobs. May be required for some older operating systems or systems with an incorrect io\_uring implementation.                                                                                                                                                                                                                                                                                                                                                                                                                                                  | `true`                      |
| `CRYOSPARC_LICENSE_SERVER_ADDR`                | Override CryoSPARC license server address. Use for systems that require access through a proxy.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                               | `https://get.cryosparc.com` |
| `CRYOSPARC_LIVE_DATA_MANAGEMENT_SCRIPT_ENABLE` | Set to `true` to run the script specified in `CRYOSPARC_LIVE_DATA_MANAGEMENT_SCRIPT_PATH` every time a CryoSPARC Live session's data management status changes. See [Live Session Data Management Tutorial](/setup-configuration-and-management/software-system-guides/cryosparc-live-session-data-management-4.7)                                                                                                                                                                                                                                                                                                                            | `false`                     |
| `CRYOSPARC_LIVE_DATA_MANAGEMENT_SCRIPT_PATH`   | The path to a bash script that runs every time a session's data management status changes: See [Live Session Data Management Tutorial](https://guide.cryosparc.com/processing-data/tutorials-and-case-studies/cryosparc-live-session-data-management). Use this to perform archive or deletion processes specific to your system (e.g., move archive data to an S3 bucket). The script is given arguments for the project UID, session UID, the datatype that changed (`micrographs`, `raw, particles`, `metadata`, or `thumbnails`), and the status of that datatype (`active`, `archiving`, `archived`, `deleted`, `deleting` or `missing`) | -                           |
| `CRYOSPARC_MONGO_CACHE_GB`                     | How much RAM to allocate for MongoDB database query cache, in GB                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              | `4`                         |
| `CRYOSPARC_MOTION_CORRECTION_CPUS_PER_GPU`     | How many CPU cores to allocate for each GPU used in Patch, Full-frame and Local Motion Correction jobs                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                        | `6`                         |
| `CRYOSPARC_MOTION_CORRECTION_RAM_MB_PER_GPU`   | How much RAM (in MB) to allocate for each GPU used in Patch, Full-frame and Local Motion Correction jobs                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      | `15000`                     |
| `CRYOSPARC_PROJECT_DIR_PREFIX`                 | The prefix to add to a project directory name on the filesystem when it is created.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           | `CS-`                       |
| `CRYOSPARC_SLACK_WEBHOOK_URL`                  | If set, CryoSPARC makes an HTTP request to this URL each time a job's status changes with some job metadata encoded in JSON                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                   | -                           |
| `CRYOSPARC_SSD_CACHE_LIFETIME_DAYS`            | Whenever a job requires SSD cache, it automatically checks for and removes files that haven't been accessed in more than the number of days specified by this variable. **Note:** Files may remain on the SSD for longer than this amount since they only get cleaned up when a job runs. Files may remain on the SSD for shorter than this amount if they are not in use by an active job and additional space is needed for caching of other particles.                                                                                                                                                                                     | `30`                        |
| `CRYOSPARC_DISABLE_EXTERNAL_REQUESTS`          | Set to `true` to prevent the application from requesting external HTTPS resources used to display information modules on the homepage.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                        | -                           |
| `REQUESTS_CA_BUNDLE`                           | May be required for connection to the CryoSPARC license verification server on systems with outdated Certificate-Authority files or that filter HTTPS requests through a proxy. Specify a path to a file or folder that contains the certificates.                                                                                                                                                                                                                                                                                                                                                                                            | -                           |
| `CRYOSPARC_HEARTBEAT_SECONDS`                  | CryoSPARC jobs running on worker nodes regularly report their status to the master command server to indicate that they are still running and active. If a job fails to report for more than this number of seconds (e.g., due to stalling, a slow network or a silent error), CryoSPARC marks the job as failed. Increase for very busy/low-resource worker nodes or slow/unreliable connections between the master and worker nodes. Increasing may reduce heartbeat-related job failures.                                                                                                                                                  | `180`                       |

**Note:** This file includes the following environment variables that are specified at installation time:

* `CRYOSPARC_LICENSE_ID`
* `CRYOSPARC_MASTER_HOSTNAME`
* `CRYOSPARC_DB_PATH`
* `CRYOSPARC_BASE_PORT`

## cryosparc\_worker/config.sh

No restart is required after changing this file.

| Variable                        | Description                                                                                                                                                                                                                                                                                                                                                                                                | Default Value |
| ------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------- |
| `CRYOSPARC_CACHE_NUM_THREADS`   | (v4.3+) Number of threads to use during caching when copying particle `.mrc` files from the project directory to the cache directory. Set to `1` to disable threading and copy files sequentially. [Guide: SSD Particle Caching in CryoSPARC](/setup-configuration-and-management/software-system-guides/tutorial-ssd-particle-caching-in-cryosparc#leveraging-multiple-threads-to-copy-particles)         | `2`           |
| `CRYOSPARC_CACHE_LOCK_STRATEGY` | (v4.5+) Distributed locking strategy to use when multiple running jobs access the SSD cache simultaneously. Set to `file` to use a file-system POSIX lock. Set to `master` to use the CryoSPARC master as a broker. Use the `master` strategy for caches on distributed file systems such as GPFS and BeeGFS where POSIX locks are disabled or unavailable.                                                | `file`        |
| `CRYOSPARC_CLIENT_TIMEOUT`      | How many seconds to wait for network requests between CryoSPARC worker/master processes before timing out with an error                                                                                                                                                                                                                                                                                    | `300`         |
| `CRYOSPARC_IMPROVED_SSD_CACHE`  | Use the new, more reliable SSD cache system (v4.4+)                                                                                                                                                                                                                                                                                                                                                        | `true`        |
| `CRYOSPARC_IO_URING`            | Set to `false` to force-disable io\_uring for fast disk read operations during jobs. May be required for some older operating systems or systems with an incorrect io\_uring implementation.                                                                                                                                                                                                               | `true`        |
| `CRYOSPARC_NO_PAGELOCK`         | By default, CryoSPARC uses the CUDA driver's `pagelocked_empty` function to allocate GPU memory. This causes CUDA-driver errors on some systems. Set this variable to `true` to use `numpy.empty`                                                                                                                                                                                                          | `false`       |
| `CRYOSPARC_SSD_PATH`            | <p>Set this variable to override the SSD cache path provided when you installed the worker. Useful if the SSD cache path is generated by your cluster as an environment variable when the job is scheduled<br><br><strong>Important</strong>: The worker must be connected with a stub SSD path for this to take effect, e.g., <code>"cache\_path": "/tmp"</code> in <code>cluster\_config.json</code></p> | -             |
| `CRYOSPARC_TIFF_IO_SHM`         | (available in CryoSPARC versions 3.2 through 4.5.3) When reading TIFF or EER movie files, CryoSPARC reads the entire file into memory before decompressing. This significantly improves performance on some networked file systems but uses more memory. Set to `false` to perform decompression directly from disk                                                                                        | `true`        |

**Note:** This file includes the following environment variables that are specified at installation time:

* `CRYOSPARC_LICENSE_ID`
* `CRYOSPARC_CUDA_PATH`
* `CRYOSPARC_USE_GPU`


# cryosparcm reference (≤v4.7)

How to use the cryosparcm utility for starting and stopping the CryoSPARC instance, checking status or logs, managing users and using CryoSPARC's command-line interface.

{% hint style="warning" %}
This page refers to CryoSPARC ≤v4.7.

For v5.0+, please see [cryosparcm reference (v5.0+)](/setup-configuration-and-management/management-and-monitoring-v5.0/cryosparcm-reference-v5.0)
{% endhint %}

## Access the command line utility, cryosparcm

The CryoSPARC master node hosts the web server and manages job resource allocation.

Workstations or master nodes with a `cryosparc_master` installation have access to `cryosparcm`, CryoSPARC's built-in [command-line](https://en.wikipedia.org/wiki/Command-line_interface) utility for all administrative, management and advanced usage tasks.

To use it, log into the machine onto which [CryoSPARC was installed](/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc). Open a Terminal running a shell (such as `bash`) and enter any of the commands described below.

```bash
cryosparcuser@emserver1:~$ cryosparcm status
----------------------------------------------------------------------------
CryoSPARC System master node installed at
/home/cryosparcuser/sw/cryosparc_master
Current cryoSPARC version: v4.3.1
----------------------------------------------------------------------------

CryoSPARC process status:

app                              RUNNING   pid 1599310, uptime 0:30:11
app_api                          RUNNING   pid 1599643, uptime 0:30:10
app_api_dev                      STOPPED   Not started
app_legacy                       STOPPED   Not started
app_legacy_dev                   STOPPED   Not started
command_core                     RUNNING   pid 1593955, uptime 0:30:26
command_rtp                      RUNNING   pid 1595620, uptime 0:30:16
command_vis                      RUNNING   pid 1594923, uptime 0:30:17
database                         RUNNING   pid 1593553, uptime 0:30:29

----------------------------------------------------------------------------
License is valid
----------------------------------------------------------------------------

global config variables:
export CRYOSPARC_LICENSE_ID="xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
export CRYOSPARC_MASTER_HOSTNAME=emserver1.internal
export CRYOSPARC_DB_PATH="/home/cryosparcuser/sw/cryosparc_db_61561"
export CRYOSPARC_BASE_PORT=61560
export CRYOSPARC_DB_CONNECTION_TIMEOUT_MS=20000
export CRYOSPARC_INSECURE=false
export CRYOSPARC_DB_ENABLE_AUTH=true
export CRYOSPARC_CLUSTER_JOB_MONITOR_INTERVAL=10
export CRYOSPARC_CLUSTER_JOB_MONITOR_MAX_RETRIES=1000000
export CRYOSPARC_PROJECT_DIR_PREFIX='CS-'
export CRYOSPARC_DEVELOP=false
export CRYOSPARC_CLICK_WRAP=true

cryosparcuser@emserver1:~$
```

{% hint style="info" %}
For systems that do not support this, navigate to the `cryosparc_master` installation directory and run `./bin/cryosparcm` instead of `cryosparcm`
{% endhint %}

## Instance Setup

### `cryosparcm update`

See [**Software Updates**](/setup-configuration-and-management/software-updates) for details.

{% content-ref url="/pages/-M7DHIJwDbQvhcKm7avb" %}
[Software Updates and Patches](/setup-configuration-and-management/software-updates)
{% endcontent-ref %}

{% hint style="warning" %}
This command can only be run by the UNIX user that owns the CryoSPARC installation directory.
{% endhint %}

### `cryosparcm patch`

Apply the latest patches available for your installed version of CryoSPARC without doing a full version update.

Specify the `--help` flag to see full usage

```
$ cryosparcm patch --help
usage: cryosparcm patch [-h] [-f] [-y] [--check] [--download] [--install]

Download and apply cryoSPARC patches. Run cryosparcm patch to automatically
install the latest patches on master and worker nodes

optional arguments:
  -h, --help   show this help message and exit
  -f, --force  install latest patch again even if already installed
  -y, --yes    confirm patch installation without prompt
  --check      check to see if a patch is available
  --download   download master and worker patches for manual installation
  --install    manually install a downloaded patch file
```

Frequently used commands:

* `cryosparcm patch`: Automatically installs the latest patches on workstations or master node and connected workers. *Not recommended for clusters: Use the `--download` and `--install` flags instead.*
* `cryosparcm patch --force`: Reinstalls the latest patches in case something went wrong with a previous attempt
* `cryosparcm patch --check`: Shows information about the latest patches without installing
* `cryosparcm patch --download`: Downloads the latest patches without installing them. Follow the resulting instructions to install the master and worker patches
* `cryosparcm patch --install`: Run this command immediately after a `--download` to install the patch on the master node.

[See also `cryosparcw patch`](/setup-configuration-and-management/management-and-monitoring-4.7/cryosparcw-4.7#cryosparcw-patch)

{% hint style="info" %}
Some patches require restarting CryoSPARC. After running `cryosparcm patch`, when prompted restart with [`cryosparcm restart`](#cryosparcm-restart).
{% endhint %}

{% hint style="warning" %}
This command can only be run by the UNIX user that owns the CryoSPARC installation directory.
{% endhint %}

### `cryosparcm cluster`

Configures cluster installation. See the [**Download and Installation**](/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc) page for details.

{% content-ref url="/pages/-M7DHIJu6M1ubOlnm8Vp" %}
[Downloading and Installing CryoSPARC](/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc)
{% endcontent-ref %}

Use with one of the following sub-commands.

#### `cryosparcm cluster example <cluster_type>`

Writes config and script template files to current working directory, `cluster_info.json` and `cluster_script.sh` respectively.

Examples are available for [Portable Batch System](https://en.wikipedia.org/wiki/Portable_Batch_System) (`cryosparcm cluster example pbs`) and [SLURM](https://slurm.schedmd.com/documentation.html) (`cryosparcm cluster example slurm`) schedulers. Other systems are similar; run one of the two `cluster example` commands and modify the output files accordingly.

#### `cryosparcm cluster dump <name>`

Outputs the existing config and script to current working directory for the cluster with the given name.

#### `cryosparcm cluster connect`

Reads `cluster_info.json` and `cluster_script.sh` from the current directory. Connects a new or updates an existing cluster configuration using the name from `cluster_info.json`.

#### `cryosparcm cluster remove <name>`

Removes a cluster configuration from the scheduler.

### `cryosparcm test`

Verifies the instance has been correctly installed by running several tests. Provides a report upon completion. For more information, see [Guide: Installation Testing with cryosparcm test](/setup-configuration-and-management/software-system-guides/guide-installation-testing-with-cryosparcm-test)

Specify the `--help` flag to see full usage

**`cryosparcm test install`**

Tests the core components of CryoSPARC (HTTP connections, licensing, workers, etc.) that are required to start running jobs and provides information on the status of the CryoSPARC instance (e.g., which version is running, whether a patch is available, etc.).

**`cryosparcm test workers <project_uid>`**

Tests workers connected to CryoSPARC to ensure they can correctly run CryoSPARC jobs by testing if the worker can launch jobs, cache particles to an SSD (if an SSD is configured), and utilize the GPU correctly.

## Instance Status and Management

The following commands are only allowed to be executed by `<cryosparcuser>` (the user that installed CryoSPARC), and can only be executed on the master node. If these conditions are not met, you may see the following error:

![Error message shown when a core command is not executed on the master node](/files/-M7DHKOR70dnethQAE2N)

You can mitigate errors temporarily by specifying the `CRYOSPARC_FORCE_HOSTNAME` variable just before calling the command:

```bash
cryosparcuser@wrong.server:~$ CRYOSPARC_FORCE_HOSTNAME="true"
cryosparcuser@wrong.server:~$ cryosparcm status
----------------------------------------------------------------------------
CryoSPARC System master node installed at
/home/cryosparcuser/cryosparc/cryosparc_master
----------------------------------------------------------------------------

cryosparcm process status:
...
```

If this error message is incorrect (the hostname specified in the error message is actually the same host, just a different identifier), you can set the hostname that the management script will use to compare by adding the variable `CRYOSPARC_HOSTNAME_CHECK` to `cryosparc_master/config.sh`. You can also set `CRYOSPARC_FORCE_HOSTNAME` in this file to permanently suppress this error.

```
File: /home/cryosparcuser/cryosparc/cryosparc_master/config.sh                          

export CRYOSPARC_LICENSE_ID="xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
export CRYOSPARC_MASTER_HOSTNAME="uoft"
export CRYOSPARC_DB_PATH="/home/cryosparcuser/cryosparc/cryosparc_master"
export CRYOSPARC_BASE_PORT=61000
export CRYOSPARC_HOSTNAME_CHECK="uoft"
```

### Available CryoSPARC Services

* `app`
* `app_api`
* `app_legacy`
* `command_core`
* `command_rtp`
* `command_vis`
* `database`

### `cryosparcm status`

Prints out current status of the CryoSPARC master system, including the status of all individual processes (`database`, `app`, `command_core`, etc). Also prints out configuration environment variables.

{% hint style="warning" %}
This command can only be run by the UNIX user that owns the CryoSPARC installation directory.
{% endhint %}

### `cryosparcm start [<service name>]`

Starts the CryoSPARC instance if stopped, including the database, the command server, the web interface, etc.

All processes start in the background, including all jobs and the web interface; the command line may be closed after running `start`.

Providing an optional service name will only start that specific service. For a list of all available services, see [Available CryoSPARC Services](#undefined).

{% hint style="warning" %}
In v4.0.0, `app_legacy` (the original web application, formerly "`webapp`") is not started by default, but can be at any time.
{% endhint %}

{% hint style="warning" %}
This command can only be run by the UNIX user that owns the CryoSPARC installation directory.
{% endhint %}

### `cryosparcm stop [<service name>]`

Stops the CryoSPARC instance if running. This will gracefully kill all the master processes, and will cause any running jobs (potentially on other nodes) to fail.

Providing an optional service name will only stop that specific service. For a list of all available services, see [Available CryoSPARC Services](#undefined).

{% hint style="warning" %}
This command can only be run by the UNIX user that owns the CryoSPARC installation directory.
{% endhint %}

### `cryosparcm restart [<service name>]`

Equivalent to running `cryosparcm stop && cryosparcm start`

Providing an optional service name will only restart that specific service. For a list of all available services, see [Available CryoSPARC Services](#undefined).

Stops the CryoSPARC instance if running. This will gracefully kill all the master processes, and will cause any running jobs (potentially on other nodes) to fail.

### `cryosparcm jobstatus`

Show a summary of how many jobs are active.

### `cryosparcm backup`

Backs up the CryoSPARC MongoDB database using `mongodump`. By default, the function creates a folder named `backup` inside the database path specified by the `CRYOSPARC_DB_PATH` environment variable in `cryosparc_master/config.sh` (e.g., `/u/cryosparcuser/cryosparc_db/backup`) and saves the backup as an `.archive` file with the current date and time in its path (e.g., `cryosparc_backup_2021_06_14_11h27.archive`) inside this folder.

{% hint style="danger" %}
Do not allow the backup to fill up the filesystem on which the database is stored. If needed, specify a custom alternative path where the backup will be written.
{% endhint %}

To change the directory the backup will be written to, specify the `--dir` flag. To change the name of the backup file, specify the `--file` flag.

{% hint style="danger" %}
CryoSPARC can be running when `cryosparcm backup` is run, but the backup will impact the performance of your running database. ([source](https://www.mongodb.com/docs/v3.6/tutorial/backup-and-restore-tools/#back-up-and-restore-with-mongodb-tools)).

Moreover, and perhaps more importantly, once CryoSPARC projects or jobs are created, deleted, or otherwise modified during or after the backup, a database restored from the resulting backup file will no longer be compatible with the modified project directories.
{% endhint %}

```bash
cryosparcm backup \
    --dir=/path/to/backups \
    --file=custombackupfile
```

### `cryosparcm restore`

Restore your database from a backup file (see [`cryosparcm backup`](#cryosparcm-backup)).

{% hint style="danger" %}
A database backup becomes outdated and incompatible with project directories as soon as CryoSPARC projects or jobs are created, deleted or modified following the the database backup. Do not restore an outdated database backup. Restoration of an outdated database backup and subsequent use with CryoSPARC is likely to corrupt CryoSPARC projects.
{% endhint %}

{% hint style="info" %}
CryoSPARC must be turned off before running this command.
{% endhint %}

{% hint style="info" %}
The database directory (`CRYOSPARC_DB_PATH` found in `cryosparc_master/config.sh`) must exist and be empty before running this command.
{% endhint %}

```
cryosparcm restore --file=/path/to/backups/backupfile
```

{% hint style="warning" %}
This command can only be run by the UNIX user that owns the CryoSPARC installation directory.
{% endhint %}

### `cryosparcm changeport <port>`

Change CryoSPARC's base port to something else. Use the `--yes` or `-y` flag to proceed without confirmation.

```
cryosparcm changeport 40000
```

{% hint style="warning" %}
This command can only be run by the UNIX user that owns the CryoSPARC installation directory.
{% endhint %}

### `cryosparcm checkdb`

Ensure the database is running with the correct host configuration

### `cryosparcm fixdbport`

Run this command after manually changing CryoSPARC's base port number to ensure the Mongo database registers the change.

### `cryosparcm licensestatus`

Check configured CryoSPARC license ID to see if there are any issues with verifying it.

### `cryosparcm maintenancemode [on|off|status]`

Stops queued jobs from running while allowing running jobs to finish. Can be used to facilitate a better user experience while CryoSPARC is undergoing maintenance, for example during restart, patch, or update. For more information, see [Guide: Maintenance Mode and Configurable User Facing Messages](/setup-configuration-and-management/software-system-guides/guide-maintenance-mode-and-configurable-user-facing-messages)

### `cryosparcm test`

Run `cryosparcm test --help` for usage information.

Test all components of a CryoSPARC instance to confirm it is working properly. For more information, see [Guide: Installation Testing with cryosparcm test](/setup-configuration-and-management/software-system-guides/guide-installation-testing-with-cryosparcm-test)

## Logs

The CryoSPARC system maintains several log files for its various processes that help with debugging any components of the CryoSPARC system that are not working.

### `cryosparcm log <process>`

Tails the output log of the master node process denoted by `<process>` which can be one of the following:

* `command_core`
* `command_rtp`
* `command_vis`
* `database`
* `app`
* `app_legacy`
* `app_api`

For example

```
cryosparcm log command_core
```

The log is live and automatically updates while the command-line remains open and new data is added to the log. To stop live updates and return to the shell, press `control C` on your keyboard, and then `q`.

To save the full log, redirect the output to a file

```
cryosparcm log command_core > command_core.log
```

To show only the last $$x$$ lines of the log, use `tail`. For example, to see the last 1000 lines of the log:

```
cryosparcm log command_core | tail -n 1000
```

### `cryosparcm filterlog <process>`

Shows the output log of the master node process denoted by `<process>` which can be one of the following:

* `command_core`
* `command_rtp`
* `command_vis`
* `database`
* `app`
* `app_legacy`
* `app_api`

For example

```
cryosparcm filterlog command_core
```

The log is live and automatically updates while the command-line remains open and new data is added to the log. To stop live updates and return to the shell, press `control C` on your keyboard.

#### Additional Arguments

To see `cryosparcm filterlog` usage, run `cryosparcm filterlog -h`:

```
Usage:
    cryosparcm filterlog SERVICE
    cryosparcm filterlog [--days|-d N] [--date|-D YYYY-MM-DD] [--max_lines|-m MAX_LINES] [--name|-n NAME] [--func|-f FUNCTION] [--level|-l LEVEL] [--tail|-t] SERVICE
Where SERVICE is one of:
    app
    app_legacy
    app_api
    command_core
    command_rtp
    command_vis
    database
Some flags not available for all services.
```

Please note that only `command_core`, `command_rtp` and `command_vis` support the following arguments with the exception of `database`, which also supports the date filter.

`--days|-d N`

* Show last N days worth of logs

`--date|-D YYYY-MM-DD`

* Show only logs from the given date

`--max_lines|-m MAX_LINES`

* the maximum number of lines to return

`--name|-n NAME`

* Show only logs with the given log name as a prefix (before dots). For example: `cryosparcm filterlog command_core -n COMMAND.SCHEDULER`

`--func|-f FUNCTION`

* Show only logs from the given log function. For example `cryosparcm filterlog command_core -f get_gpu_info_run`

`--level|-l LEVEL`

* Show only logs with the given level. For example: `cryosparcm filterlog command_core -l ERROR`

`--tail|-t`

* Will tail the log with the filters applied. To stop live updates and return to the shell, press `control C.`

### `cryosparcm joblog PX JXX`

Shows a live output log for job `JXX` in project `PX`. Includes the standard input and error from the python process for the job.

For example, to show the output of Job 123 in Project 3, run the following

```
cryosparcm joblog P3 J123
```

{% hint style="info" %}
`joblog` shows the full `stdout` for the job, which is more comprehensive than the job log in the web interface and is more helpful for debugging.
{% endhint %}

To stop logging, save the full log or view only the last few log lines, see instructions for [`cryosparcm log`](#cryosparcm-log-less-than-process-greater-than)

### `cryosparcm eventlog PX JXX`

Prints the text of the processing log for job `JXX` in project `PX` to stdout.

For example, to write text components of the processing log for CryoSPARC job J123 in project P3 to a file in the current working directory, run:

`cryosparcm eventlog P3 J123 > P3_J123_events.log`

### `cryosparcm snaplogs`

Compresses all `.log` files in the `cryosparc_master/run` folder into a `.tgz` file inside `cryosparc_master/run`.

### `cryosparcm errorreport`

Run `cryosparcm errorreport --help` for usage information.

Create a CryoSPARC error report which includes diagnostic information and CryoSPARC instance logs. For more information, see [Guide: Download Error Reports](/setup-configuration-and-management/software-system-guides/guide-download-error-reports)

## User Management

### `cryosparcm listusers`

Outputs a table of users registered with the system, including their names, email addresses and admin status.

### `cryosparcm createuser`

Creates a new user account for log in through the web interface. Full use:

```bash
cryosparcm createuser \
    --email "<email address>" \
    --username "<login username>" \
    --firstname "<given name>"\
    --lastname "<surname>" \
    [--password "<new password>"]
```

{% hint style="info" %}
If the`--password` argument is not specified, a silent input prompt is provided.
{% endhint %}

{% hint style="warning" %}
This command can only be run by the UNIX user that owns the CryoSPARC installation directory.
{% endhint %}

### `cryosparcm resetpassword`

Resets the password for indicated user with the new `<password>` provided. Full use:

```bash
cryosparcm resetpassword \
    --email "<email address>" \
    [--password "<new password>"]
```

{% hint style="info" %}
If the `--password` argument is not specified, a silent input prompt is provided.
{% endhint %}

{% hint style="warning" %}
This command can only be run by the UNIX user that owns the CryoSPARC installation directory.
{% endhint %}

### `cryosparcm updateuser`

Updates a user profile to change the user name or set admin privileges. Full use:

```bash
cryosparcm updateuser \
  --email "<email associated with cryoSPARC account>" \
  --username "<new username for that user>" \
  --firstname "<new first name(s) for that user>" \
  --lastname "<new last name(s) for that user>" \
  --admin "<true|false>" \
  [--password "<corresponding cryoSPARC password>"]
```

{% hint style="info" %}
If the `--password` argument is not specified, a silent input prompt is provided.
{% endhint %}

`--email` is required and can *not* be changed with this command.

`--username`, `--firstname`, `--lastname` are used for display and do not otherwise affect CryoSPARC function.

Other than for the first-created user account, new users do not have administrative privileges applied by default. Following creation of the first-created user account, other accounts can also be created through the user interface if preferred:

{% content-ref url="/pages/-MNecrUQzW1CYEnJCepj" %}
[Guide: User Management](/setup-configuration-and-management/software-system-guides/tutorial-user-management)
{% endcontent-ref %}

{% hint style="warning" %}
This command can only be run by the UNIX user that owns the CryoSPARC installation directory.
{% endhint %}

## Command-Line Interface

### `cryosparcm env`

Prints a list of environment variables for use in a command-line shell to replicate the exact environment used to when running CryoSPARC processes.

Run this command with `eval` to define the variables output by the `env` command.

```
eval $(cryosparcm env)
```

This may be used to, for example, run the [**Python**](https://www.python.org/) distribution that ships with CryoSPARC.

### `cryosparcm cli`

Runs a command using with cryoSPARC's command-line interface

```
cryosparcm cli <command>
```

See the full command reference

{% content-ref url="/pages/-M7DHIJzq3PnUov4PO7G" %}
[cryosparcm cli reference (≤v4.7)](/setup-configuration-and-management/management-and-monitoring-4.7/cli-4.7)
{% endcontent-ref %}

### `cryosparcm icli`

Runs an an interactive cryosparc shell that connects to the master processes. Use this to interactively run Python commands in the same environment as CryoSPARC.

![](/files/-M7DHKOWSMop88e2BjWb)

{% content-ref url="/pages/-M7DHIJzq3PnUov4PO7G" %}
[cryosparcm cli reference (≤v4.7)](/setup-configuration-and-management/management-and-monitoring-4.7/cli-4.7)
{% endcontent-ref %}

### `cryosparcm rtpcli`

Similar to [`cryosparcm cli`](/setup-configuration-and-management/management-and-monitoring-4.7/cryosparcw-4.7#cryosparcm-cli), but for running interactive commands with CryoSPARC Live.

**Additional details coming soon.**

See our guide on Managing a CryoSPARC Live Session from the CLI:

{% embed url="<https://guide.cryosparc.com/live/how-to-access-cryosparc-live/managing-a-cryosparc-live-session-from-the-cli#setup>" %}
<https://guide.cryosparc.com/live/how-to-access-cryosparc-live/managing-a-cryosparc-live-session-from-the-cli#setup>
{% endembed %}

### `cryosparcm downloadtest`

Download test data (subset of 20 movies from the EMPIAR-10025 dataset) to the current working directory for use with the [**T20S Introductory Tutorial**](/guides-for-v3/cryo-em-data-processing-in-cryosparc-introductory-tutorial#t-20-s-tutorial).

### `cryosparcm help`

Prints a help message

### `cryosparcm mongo`

Starts a [`mongo` **shell**](https://docs.mongodb.com/manual/mongo/) for CryoSPARC's local mongoDB instance.

### `cryosparcm call <command>`

Execute a shell command in CryoSPARC's shell environment, such as Python. Equivalent to calling:

`eval $(cryosparcm env)` followed by another shell command.


# cryosparcm cli reference (≤v4.7)

How to use CryoSPARC's low-level command-line interface.

{% hint style="warning" %}
This page refers to CryoSPARC ≤v4.7.

For v5.0+, please see [cryosparcm cli reference (v5.0+)](/setup-configuration-and-management/management-and-monitoring-v5.0/cryosparcm-cli-reference-v5.0)
{% endhint %}

Call methods described in this module directly as arguments to `cryosparcm cli` or referenced from the `cli` object in `cryosparcm icli`.

* **cli example:**

```bash
cryosparcm cli "enqueue_job(project_uid='P3', job_uid='J42', lane='cryoem1')"
```

* **icli example:**

```bash
$ cryosparcm icli

connecting to cryoem5:61002 ...
cli, rtp, db, gfs and tools ready to use

In [1]: cli.enqueue_job(project_uid='P3', job_uid='J42', lane='cryoem1')

In [2]:
```

### `add_project_user_access(project_uid: str, requester_user_id: str, add_user_id: str)`

Allows owner of a project (or admin) to grant access to another user to view and edit the project

* **Parameters:**
  * **project\_uid** (*str*) -- uid of the project to grant access to
  * **requester\_user\_id** (*str*) -- \_id of the user requesting a new user to be added to the project
  * **add\_user\_id** (*str*) -- uid of the user to add to the project
* **Raises:** AssertionError

### `add_scheduler_lane(name: str, lanetype: str, title: str | None = None, desc: str = '')`

Adds a new lane to the master scheduler

* **Parameters:**
  * **name** (*str*) -- name of the lane
  * **lanetype** (*str*) -- type of lane ("cluster" or "node")
  * **title** (*str*) -- optional title of the lane
  * **desc** (*str*) -- optional description of the lane

### `add_scheduler_target_cluster(name, worker_bin_path, script_tpl, send_cmd_tpl='{{ command }}', qsub_cmd_tpl='qsub {{ script_path_abs }}', qstat_cmd_tpl='qstat -as {{ cluster_job_id }}', qstat_code_cmd_tpl=None, qdel_cmd_tpl='qdel {{ cluster_job_id }}', qinfo_cmd_tpl='qstat -q', transfer_cmd_tpl='cp {{ src_path }} {{ dest_path }}', cache_path=None, cache_quota_mb=None, cache_reserve_mb=10000, title=None, desc=None)`

Add a cluster to the master scheduler

* **Parameters:**
  * **name** (*str*) -- name of cluster
  * **worker\_bin\_path** (*str*) -- absolute path to 'cryosparc\_package/cryosparc\_worker/bin' on the cluster
  * **script\_tpl** (*str*) -- script template string
  * **send\_cmd\_tpl** (*str*) -- send command template string
  * **qsub\_cmd\_tpl** (*str*) -- queue submit command template string
  * **qstat\_cmd\_tpl** (*str*) -- queue stat command template string
  * **qdel\_cmd\_tpl** (*str*) -- queue delete command template string
  * **qinfo\_cmd\_tpl** (*str*) -- queue info command template string
  * **transfer\_cmd\_tpl** (*str*) -- transfer command template string (currently unused)
  * **cache\_path** (*str*) -- path on SSD that can be used for the cryosparc cache
  * **cache\_quota\_mb** (*int*) -- the max size (in MB) to use for the cache on the SSD
  * **cache\_reserve\_mb** (*int*) -- size (in MB) to set aside for uses other than CryoSPARC cache
  * **title** (*str*) -- an optional title to give to the cluster
  * **desc** (*str*) -- an optional description of the cluster
* **Returns:** configuration parameters for new cluster
* **Return type:** dict

### `add_scheduler_target_node(hostname, ssh_str, worker_bin_path, num_cpus, cuda_devs, ram_mb, has_ssd, cache_path=None, cache_quota=None, cache_reserve=10000, monitor_port=None, gpu_info=None, lane='default', title=None, desc=None)`

Adds a worker node to the master scheduler

* **Parameters:**
  * **hostname** (*str*) -- the hostname of the target worker node
  * **ssh\_str** (*str*) -- the ssh connection string of the worker node
  * **worker\_bin\_path** (*str*) -- the absolute path to 'cryosparc\_package/cryosparc\_worker/bin' on the worker node
  * **num\_cpus** (*int*) -- total number of CPU threads available on the worker node
  * **cuda\_devs** (*list*) -- total number of cuda-capable devices (GPUs) on the worker node, represented as a list (i.e. range(4))
  * **ram\_mb** (*float*) -- total available physical ram in MB
  * **has\_ssd** (*bool*) -- if an ssd is available or not
  * **cache\_path** (*str*) -- path on SSD that can be used for the cryosparc cache
  * **cache\_quota** (*int*) -- the max size (in MB) to use for the cache on the SSD
  * **cache\_reserve** (*int*) -- size (in MB) to initially reserve for the cache on the SSD
  * **gpu\_info** (*list*) -- compiled GPU information as computed by get\_gpu\_info
  * **lane** (*str*) -- the scheduler lane to add the worker node to
  * **title** (*str*) -- an optional title to give to the worker node
  * **desc** (*str*) -- an optional description of the worker node
* **Returns:** configuration parameters for new worker node
* **Return type:** dict
* **Raises:** AssertionError

### `add_tag_to_job(project_uid: str, job_uid: str, tag_uid: str)`

Tag the given job with the given tag UID

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **job\_uid** (*str*) -- target job UID, e.g., "J42"
  * **tag\_uid** (*str*) -- target tag UID, e.g., "T1"
* **Returns:** contains modified jobs count
* **Return type:** dict

### `add_tag_to_project(project_uid: str, tag_uid: str)`

Tag the given project with the given tag

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **tag\_uid** (*str*) -- target tag UID, e.g., "T1"
* **Returns:** contains modified documents count
* **Return type:** dict

### `add_tag_to_session(project_uid: str, session_uid: str, tag_uid: str)`

Tag the given session with the given tag

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **session\_uid** (*str*) -- target session UID, e.g., "S1"
  * **tag\_uid** (*str*) -- target tag UID, e.g., "T1"
* **Returns:** contains modified sessions count
* **Return type:** dict

### `add_tag_to_workspace(project_uid: str, workspace_uid: str, tag_uid: str)`

Tag the given workspace with the given tag

* **Parameters:**
  * **project\_uid** (*\_type\_*) -- target project UID, e.g., "P3"
  * **workspace\_uid** (*str*) -- target workspace UID, e.g. "W1"
  * **tag\_uid** (*str*) -- target tag UID, e.g., "T1"
* **Returns:** contains modified workspaces count
* **Return type:** dict

### `admin_user_exists()`

Returns True if there exists at least one user with admin privileges, False otherwise

* **Returns:** Whether an admin exists
* **Return type:** bool

### `archive_project(project_uid: str)`

Archive given project. This means that the project can no longer be modified and jobs cannot be created or run. Once archived, the project directory may be safely moved to long-term storage.

* **Parameters:** **project\_uid** (*str*) -- target project UID, e.g., "P3"

### `attach_project(owner_user_id: str, abs_path_export_project_dir: str)`

Attach a project by importing it from the given path and writing a new lockfile. May only run this on previously-detached projects.

* **Parameters:**
  * **owner\_user\_id** (*str*) -- user account ID performing this opertation that will take ownership of imported project.
  * **abs\_path\_export\_project\_dir** (*str*) -- absolute path to directory of CryoSPARC project to attach
* **Returns:** new project UID of attached project
* **Return type:** str

### `calculate_intermediate_results_size(project_uid: str, job_uid: str, always_keep_final: bool = True, use_prt=True)`

Find intermediate results, calculate intermediate results total size, save it to the job doc, all in a PostResponseThread

### `check_project_exists(project_uid: str)`

Check that the given project exists and has not been deleted

* **Parameters:** **project\_uid** (*str*) -- unique ID of target project, e.g., "P3"
* **Returns:** True if the project exists and hasn't been deleted
* **Return type:** bool

### `check_workspace_exists(project_uid: str, workspace_uid: str)`

Returns True if target workspace exists.

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **workspace\_uid** (*str*) -- target workspace UID, e.g., "W4"
* **Returns:** True if the workspace exists, False otherwise
* **Return type:** bool

### `cleanup_data(project_uid: str, workspace_uid: str | None = None, delete_non_final: bool = False, delete_statuses: List[str] = [], clear_non_final: bool = False, clear_sections: List[str] = [], clear_types: List[str] = [], clear_statuses: List[str] = [])`

Cleanup project or workspace data, clearing/deleting jobs based on final result status, sections, types, or status

### `cleanup_jobs(project_uid: str, job_uids_to_delete: List[str], job_uids_to_clear: List[str])`

Cleanup jobs helper for cleanup\_data

### `clear_intermediate_results(project_uid: str, workspace_uid: str | None = None, job_uid: str | None = None, always_keep_final: bool | None = True)`

Asynchornously remove intermediate results from the given project or job

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **workspace\_uid** (*str*\*,\* *optional*) -- clear intermediate results for all jobs in a workspace
  * **job\_uid** (*str*\*,\* *optional*) -- target job UID, e.g., "J42", defaults to None
  * **always\_keep\_final** (*bool*\*,\* *optional*) -- not used, defaults to True

### `clear_job(project_uid: str, job_uid: str, nofail=False, isdelete=False)`

Clear a job to get it back to building state (do not clear params or inputs)

* **Parameters:**
  * **project\_uid** (*str*) -- uid of the project that contains the job to clear
  * **job\_uid** (*str*) -- uid of the job to clear
  * **nofail** (*bool*) -- If True, ignores errors that occur while job is clearing, defaults to False
  * **isdelete** (*bool*) -- Set to True when the job is about to be deleted, defaults to False
* **Raises:** AssertionError

### `clear_jobs_by_section(project_uid, workspace_uid=None, sections=[])`

Clear all jobs from a project belonging to specified sections in the job register, optionally scoped to a workspace

### `clear_jobs_by_type(project_uid, workspace_uid=None, types=[])`

Clear all jobs from a project belonging to specified types, optionally scoped to a workspace

### `clone_job(project_uid: str, workspace_uid: str | None, job_uid: str, created_by_user_id: str, created_by_job_uid: str | None = None)`

Creates a new job as a clone of the provided job

* **Parameters:**
  * **project\_uid** (*str*) -- project UID
  * **workspace\_uid** (*str* *|* *None*) -- uid of the workspace, may be empty
  * **job\_uid** (*str*) -- uid of the job to copy
  * **created\_by\_user\_id** (*str*) -- the id of the user creating the clone
  * **created\_by\_job\_uid** (*str*\*,\* *optional*) -- uid of the job creating the clone, defaults to None
* **Returns:** the job uid of the newly created clone
* **Return type:** str

### `clone_job_chain(project_uid: str, start_job_uid: str, end_job_uid: str, created_by_user_id: str, workspace_uid: str | None = None, new_workspace_title: str | None = None)`

Clone jobs that directly descend from the given start job UID up to the given end job UID. Returns a dict with information about cloned jobs, or None if nothing was cloned

* **Parameters:**
  * **project\_uid** (*str*) -- project UID where jobs are located, e.g., "P3"
  * **start\_job\_uid** (*str*) -- starting anscestor job UID
  * **end\_job\_uid** (*str*) -- ending descendant job UID
  * **created\_by\_user\_id** (*str*) -- ID of user performing this operation
  * **workspace\_uid** (*str*\*,\* *optional*) -- uid of workspace to clone jobs into, defaults to None
  * **new\_workspace\_title** (*str*\*,\* *optional*) -- Title of new workspace to create if a uid is not provided, defaults to None
* **Returns:** dictionary with information about created jobs and workspace or None
* **Return type:** dict | None

### `clone_jobs(project_uid: str, job_uids: List[str], created_by_user_id, workspace_uid=None, new_workspace_title=None)`

Clone the given list of jobs. If any jobs are related, it will try to re-create the input connections between the cloned jobs (but maintain the same connections to jobs that were not cloned)

* **Parameters:**
  * **project\_uid** (*str*) -- target project uid (e.g., "P3")
  * **job\_uids** (*list*) -- List of job UIDs to delete (e.g., `["J1", "J2", "J3"]`)
  * **created\_by\_user\_id** (*str*) -- ID of user performing this operation
  * **workspace\_uid** (*str*\*,\* *optional*) -- uid of workspace to clone jobs into, defaults to None
  * **new\_workspace\_title** (*str*\*,\* *optional*) -- Title of new workspace to create if one is not provided, defaults to None
* **Returns:** dictionary with information about created jobs and workspace
* **Return type:** dict

### `compile_extensive_validation_benchmark_data(instance_information: dict, job_timings: dict)`

Compile extensive validation benchmark data for upload

### `create_backup(backup_dir: str, backup_file: str)`

Create a database backup in the given directory with the given file name

* **Parameters:**
  * **backup\_dir** (*str*) -- directory to create backup in
  * **backup\_file** (*str*) -- filename of backup file

### `create_empty_project(owner_user_id: str, project_container_dir: str, title: str, desc: str | None = None, project_dir: str | None = None, export=True, hidden=False)`

Creates a new project and project directory and creates a new document in the project collection

#### `NOTE`

project\_container\_dir is an absolute path that the user guarantees is available everywhere.

* **Parameters:**
  * **owner\_user\_id** (*str*) -- the \_id of the user that owns the new project (which is also the user that requests to create the project)
  * **project\_container\_dir** (*str*) -- the path to the "root" directory in which to create the new project directory
  * **title** (*str*\*,\* *optional*) -- the title of the new project to create, defaults to None
  * **desc** (*str*\*,\* *optional*) -- the description of the new project to create, defaults to None
  * **export** (*bool*\*,\* *optional*) -- if True, outputs project details to disk for exporting to another instance, defaults to True
  * **hidden** (*bool*\*,\* *optional*) -- if True, outputs project details to disk for exporting to another instance, defaults to False
* **Returns:** the new uid of the project that was created
* **Return type:** str

### `create_empty_workspace(project_uid, created_by_user_id, created_by_job_uid=None, title=None, desc=None, export=True)`

Add a new empty workspace to the given project.

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **created\_by\_user\_id** (*str*) -- User \_id that creates this workspace
  * **created\_by\_job\_uid** (*str*\*,\* *optional*) -- Job UID that creates this workspace, defaults to None
  * **title** (*str*\*,\* *optional*) -- Workspace title,, defaults to None
  * **desc** (*str*\*,\* *optional*) -- Workspace description, defaults to None
  * **export** (*bool*\*,\* *optional*) -- If True, dumps workspace contents to disk for exporting to other CryoSPARC instances, defaults to True
* **Returns:** UID of new workspace, e.g., "W4"
* **Return type:** str

### `create_new_job(job_type, project_uid, workspace_uid, created_by_user_id, created_by_job_uid=None, title=None, desc=None, do_layout=True, dry_run=False, enable_bench=False, priority=None)`

Create a new job

#### `NOTE`

This request runs on its own thread, but locks on the "jobs" lock, so only one create can happen at a time.This means that the builder functions below can take as much time as they like to do stuff in their own thread (releasing GIL like IO or numpy etc) and command can still respond to other requests that don't require this lock.

* **Parameters:**
  * **job\_type** (*str*) -- the name of the job to create
  * **project\_uid** (*str*) -- uid of the project
  * **workspace\_uid** (*str*) -- uid of the workspace
  * **created\_by\_user\_id** (*str*) -- the id of the user creating the job
  * **created\_by\_job\_uid** (*str*) -- uid of the job that created this job
  * **title** (*str*) -- optional title of the job
  * **desc** (*str*) -- optional description of the job
  * **do\_layout** (*bool*) -- specifies if the job tree layout should be recalculated, defaults to True
  * **enable\_bench** (*bool*) -- should performance be measured for this job (benchmarking)
* **Returns:** uid of the newly created job
* **Return type:** str
* **Raises:** AssertionError

### `create_tag(title: str, type: str, created_by_user_id: str, colour: str | None = None, description: str | None = None)`

Create a new tag

* **Parameters:**
  * **title** (*str*) -- tag name
  * **type** (*str*) -- tag type such as "general", "project", "workspace", "session" or "job"
  * **created\_by\_user\_id** (*str*) -- user account ID performing this action
  * **colour** (*str*\*,\* *optional*) -- tag colour, may be "black", "red", green", etc., defaults to None
  * **description** (*str*\*,\* *optional*) -- detailed tag description, defaults to None
* **Returns:** created tag document
* **Return type:** dict

### `create_user(created_by_user_id: str, email: str, password: str, username: str, first_name: str, last_name: str, admin: bool = False)`

Creates a new CryoSPARC user account

* **Parameters:**
  * **created\_by\_user\_id** (*str*) -- identity of user who is creating the new user
  * **email** (*str*) -- email of user to create
  * **password** (*str*) -- password of user to create
  * **admin** (*bool*\*,\* *optional*) -- specifies if user should have an administrator role or not, defaults to False
* **Returns:** mongo \_id of user that was created
* **Return type:** str

### `delete_detached_project(project_uid: str)`

Deletes large database entries such as streamlogs and gridFS for a detached project and marks the project and its jobs as deleted

### `delete_job(project_uid: str, job_uid: str, relayout=True, nofail=False, force=False)`

Deletes a job after killing it (if running), clearing it, setting the "deleted" property, and recalculating the tree layout

* **Parameters:**
  * **project\_uid** (*str*) -- uid of the project that contains the job to delete
  * **job\_uid** (*str*) -- uid of the job to delete
  * **relayout** (*bool*) -- specifies if the tree layout should be recalculated or not

### `delete_jobs(project_job_uids: List[Tuple[str, str]], relayout=True, nofail=False, force=False)`

Delete the given project-job UID combinations, provided in the following format:

```python
[("PX", "JY"), ...]
```

Or the following is also valid:

```python
[["PX", "JY"], ...]
```

Where PX is a project UID and JY is a job UID in that project

* **Parameters:**
  * **project\_job\_uids** (*list*\*\[**tuple**\[**str**,\* *str*\*]\*\*]\*) -- project-job UIDs to delete
  * **relayout** (*bool*\*,\* *optional*) -- whether to recompute the layout of the job tree, defaults to True
  * **nofail** (*bool*\*,\* *optional*) -- if True, ignore errors when deleting, defaults to False

### `delete_project(project_uid: str, request_user_id: str, jobs_to_delete: list | None = None, workspaces_to_delete: list | None = None)`

Iterate through each job and workspace associated with the project to be deleted and delete both, and then disable the project.

* **Parameters:**
  * **project\_uid** (*str*) -- uid of the project to be deleted
  * **request\_user\_id** (*str*) -- \_id of user requesting the project to be deleted
  * **jobs\_to\_delete** (*list*) -- list of all jobs to be fully deleted
  * **workspaces\_to\_delete** (*list*) -- list of all workspaces that will be deleted by association
* **Returns:** string confirmation that the workspace has been deleted
* **Return type:** str
* **Raises:** AssertionError -- if for some reason a workspace isn't deleted, this will be thrown

### `delete_project_user_access(project_uid: str, requester_user_id: str, delete_user_id: str)`

Removes a user's access from a project.

* **Parameters:**
  * **project\_uid** (*str*) -- uid of the project to revoke access to
  * **requester\_user\_id** (*str*) -- \_id of the user requesting a user to be removed from the project
  * **delete\_user\_id** (*str*) -- uid of the user to remove from the project
* **Raises:** AssertionError

### `delete_user(email: str, requesting_user_email: str, requesting_user_password: str)`

Remove a user from the CryoSPARC. Only administrators may do this

* **Parameters:**
  * **email** (*str*) -- user's email
  * **requesting\_user\_email** (*str*) -- your CryoSPARC login email
  * **requesting\_user\_password** (*str*) -- your CryoSPARC password
* **Returns:** confirmation message
* **Return type:** str

### `delete_workspace(project_uid: str, workspace_uid: str, request_user_id: str, jobs_to_delete_inside_one_workspace: list | None = None, jobs_to_update_inside_multiple_workspaces: list | None = None)`

Asynchronously iterate through each job associated with the workspace to be deleted and either delete or update the job, and then disables the workspace

* **Parameters:**
  * **project\_uid** (*str*) -- uid of the project containing the workspace to be deleted
  * **workspace\_uid** (*str*) -- uid of the workspace to be deleted
  * **request\_user\_id** (*str*) -- uid of the user that is requesting the workspace to be deleted
  * **jobs\_to\_delete\_inside\_one\_workspace** (*list*) -- list of all jobs to be fully deleted
  * **jobs\_to\_update\_inside\_multiple\_workspaces** (*list*) -- list of all jobs to be updated so that the workspace uid is removed from its list of "workspace\_uids"
* **Returns:** string confirmation that the workspace has been deleted
* **Return type:** str

### `detach_project(project_uid: str)`

Detach a project by exporting, removing lockfile, then setting its detached property. This hides the project from the interface and allows other instances to attach and run this project.

* **Parameters:** **project\_uid** (*str*) -- target project UID to detach, e.g., "P3"

### `do_job(job_type: str, puid='P1', wuid='W1', uuid='devuser', params={}, input_group_connects={})`

Create and run a job on the "default" node

* **Parameters:**
  * **job\_type** (*str*) -- type of job, e.g, "abinit"
  * **puid** (*str*\*,\* *optional*) -- project UID to create job in, defaults to 'P1'
  * **wuid** (*str*\*,\* *optional*) -- workspace UID to create job in, defaults to 'W1'
  * **uuid** (*str*\*,\* *optional*) -- User ID performing this action, defaults to 'devuser'
  * **params** (*str*\*,\* *optional*) -- parameter overrides for the job, defaults to {}
  * **input\_group\_connects** (*dict*\*,\* *optional*) -- input group connections dictionary, where each key is the input group name and each value is the parent output identifier, e.g., `{"particles": "J1.particles"}`, defaults to {}
* **Returns:** created job UID
* **Return type:** str

### `dump_license_validation_results()`

Get a summary of license validation checks

* **Returns:** Text description of validation checks separated by newlines
* **Return type:** str

### `enqueue_job(project_uid: str, job_uid: str, lane: str | None = None, user_id: str | None = None, hostname: str | None = None, gpus: List[int] | typing_extensions.Literal[False] = False, no_check_inputs_ready: bool = False)`

Add the job in the given project to the queue for the given worker lane (default lane if not specified)

* **Parameters:**
  * **project\_uid** (*str*) -- uid of the project containing the job to queue
  * **job\_uid** (*str*) -- job uid to queue
  * **lane** (*str*\*,\* *optional*) -- name of the worker lane onto which to queue
  * **hostname** (*str*\*,\* *optional*) -- the hostname of the target worker node
  * **gpus** (*list*\*,\* *optional*) -- list of GPU indexes on which to queue the given job, defaults to False
  * **no\_check\_inputs\_ready** (*bool*\*,\* *optional*) -- if True, forgoes checking whether inputs are ready (not recommended), defaults to False
* **Returns:** job status (e.g., 'launched' if success)
* **Return type:** str
* **Example:**

```python
enqueue_job('P3', 'J42', lane='cryoem1', user_id='62b64b77632103020e4e30a7')
```

### `export_job(project_uid: str, job_uid: str)`

Start export for the given job into the project's exports directory

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **job\_uid** (*str*) -- target job UID to export, e.g., "J42"

### `export_output_result_group(project_uid: str, job_uid: str, output_result_group_name: str, result_names: List[str] | None = None)`

Start export of given output result group. Optionally specify which output slots to select for export

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **job\_uid** (*str*) -- target job UID, e.g,. "J42"
  * **output\_result\_group\_name** (*str*) -- target output, e.g., "particles"
  * **result\_names** (*list*\*,\* *optional*) -- target result slots list, e.g., \["blob", "location"], exports all if unspecified, defaults to None

### `export_project(project_uid: str, override=False)`

Ensure the given project is ready for transfer to another instance. Call `detach_project` once this is complete.

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **override** (*bool*\*,\* *optional*) -- force run even if recently exported, defaults to False

### `generate_new_instance_uid(force_takeover_projects=False)`

Generates a new uid for the CryoSPARC instance.

* **Parameters:** **force\_takeover\_projects** (*boolean*\*,\* *optional*) -- If True, take overwrite existing lockfiles. If False, only creates lockfile in projects that don't already have one. Defaults to False
* **Returns:** New instance UID
* **Return type:** str

### `get_all_tags()`

Get a list of available tag documents

* **Returns:** list of tags
* **Return type:** list

### `get_base_template_args(project_uid, job_uid, target)`

Returns template arguments for cluster commands including: project\_uid job\_uid job\_creator cryosparc\_username project\_dir\_abs job\_dir\_abs job\_log\_path\_abs job\_type

### `get_cryosparc_log(service_name, days=7, date=None, max_lines=None, log_name='', func_name='', level='')`

Get cryosparc service logs, filterable by date, name, function, and level

* **Parameters:**
  * **service\_name** (*str*) -- target service name
  * **days** (*int*\*,\* *optional*) -- number of days of logs to retrive, defaults to 7
  * **date** (*str*\*,\* *optional*) -- retrieve logs from a specific day formatted as "YYYY-MM-DD", defaults to None
  * **max\_lines** (*int*\*,\* *optional*) -- maximum number of lines to retrieve, defaults to None
  * **log\_name** (*str*\*,\* *optional*) -- name of internal logger type such as ", defaults to ""
  * **func\_name** (*str*\*,\* *optional*) -- name of Python function producing logs, defaults to ""
  * **level** (*str*\*,\* *optional*) -- log severity level such as "INFO", "WARNING" or "ERROR", defaults to ""
* **Returns:** Filtered log
* **Return type:** str

### `get_default_job_priority(created_by_user_id: str)`

Get the default job priority for jobs queued by the given user ID

* **Parameters:** **created\_by\_user\_id** (*str*) -- target user account \_id
* **Returns:** the integer priority, with 0 being the highest priority
* **Return type:** int

### `get_gpu_info()`

Asynchronously update the available GPU information on all connected nodes

### `get_instance_default_job_priority()`

Default priority for jobs queued by users without an explicit priority set in Admin > User Management

* **Returns:** Default priority config variable, 0 if never set
* **Return type:** int

### `get_job(project_uid: str, job_uid: str, \*args, \*\*kwargs)`

Get a job object with optionally only a few fields specified in `*args`

* **Parameters:**
  * **project\_uid** (*str*) -- project UID
  * **job\_uid** (*str*) -- uid of job
* **Returns:** mongo result set
* **Return type:** dict
* **Example:**

```python
get_job('P1', 'J1', 'status', 'job_type', 'project_uid')
```

### `get_job_chain(project_uid: str, start_job_uid: str, end_job_uid: str)`

Get a list of jobs all jobs that are descendants of the start job and ancestors of the end job.

* **Parameters:**
  * **project\_uid** (*str*) -- project UID where jobs are located, e.g., "P3"
  * **start\_job\_uid** (*str*) -- starting ascenstor job UID
  * **end\_job\_uid** (*str*) -- ending descendant job UID
* **Returns:** list of jobs in the job chain
* **Return type:** str

### `get_job_creator_and_username(job_uid, user=None)`

Returns job creator and username strings for cluster job submission

### `get_job_dir_abs(project_uid: str, job_uid: str)`

Get the path to the given job's directory

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **job\_uid** (*str*) -- target job UID, e.g., "J42"
* **Returns:** absolute path to job log directory
* **Return type:** str

### `get_job_log(project_uid: str, job_uid: str)`

Get the full contents of the given job's standard output log

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **job\_uid** (*str*) -- target job UID, e.g., "J42"
* **Returns:** job log contents
* **Return type:** str

### `get_job_log_path_abs(project_uid: str, job_uid: str)`

Get the path to the given job's standard output log

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **job\_uid** (*str*) -- target job UID, e.g., "J42"
* **Returns:** absolute path to job log file
* **Return type:** str

### `get_job_output_min_fields(project_uid: str, job_uid: str, output: str)`

Get the minimum expected fields description for the given output name.

### `get_job_queue_message(project_uid: str, job_uid: str)`

Get message stating why a job is queued but not running

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **job\_uid** (*str*) -- job UID, e.g., "J42"
* **Returns:** queue message
* **Return type:** str

### `get_job_result(project_uid: str, source_result: str)`

Get the `output_results` item that matches this source result

* **Code-block::** python

  source\_result = JXX.output\_group\_name.result\_name
* **Parameters:**
  * **project\_uid** (*str*) -- id of the project
  * **source\_result** (*str*) -- Source result
* **Returns:** the result details from the database
* **Return type:** dict

### `get_job_sections()`

Retrieve available jobs, origanized by section

* **Returns:** job sections
* **Return type:** list\[dict\[str, Any]]

### `get_job_status(project_uid: str, job_uid: str)`

Get the status of the given job

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **job\_uid** (*str*) -- target job UID, e.g., "J42"
* **Returns:** job status such as "building", "queued" or "completed"
* **Return type:** str

### `get_job_streamlog(project_uid: str, job_uid: str)`

Get a list of dictionaries representing the given job's event log

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **job\_uid** (*str*) -- target job UID, e.g., "J42"
* **Returns:** event log contents
* **Return type:** str

### `get_job_symlinks(project_uid: str, job_uid: str)`

Get a list of symbolic links in the given job directory

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g, "P3"
  * **job\_uid** (*str*) -- target job UID, e.g., "J42"
* **Returns:** List of symlinks paths, either absolute or relative to job directory
* **Return type:** list

### `get_jobs_by_section(project_uid, workspace_uid=None, sections=[], fields=[])`

Get all jobs in a project belonging to the specified sections in the job register, optionally scoped to a workspace

### `get_jobs_by_status(status: str, count_only=False)`

Get the list of jobs with the given status (or count if `count_only` is True)

* **Parameters:**
  * **status** (*str*) -- Possible job status such as "building", "queued" or "completed"
  * **count\_only** (*bool*\*,\* *optional*) -- If True, only return the integer count, defaults to False
* **Returns:** List of job documents or count
* **Return type:** list | int

### `get_jobs_by_type(project_uid, workspace_uid=None, types=[], fields=[])`

Get all jobs matching the given types, optionally scoped to a workspace

### `get_jobs_size_by_section(project_uid: str, workspace_uid: str = None, sections: List[str] = [])`

Gets the total job size for jobs belonging to the input sections of the job register

### `get_last_backup_complete_activity()`

Get details about most recent database backup

* **Returns:** activity document or {} if never backed up
* **Return type:** dict

### `get_maintenance_mode()`

Get maintenance mode status.

* **Returns:** True if set, False otherwise
* **Return type:** bool

### `get_non_final_jobs(project_uid: str, workspace_uid: str = None, fields: List[str] = [])`

Gets a list of job uids of all jobs not marked as a final result or ancestor of final result

### `get_non_final_jobs_size(project_uid: str, workspace_uid: str = None)`

Gets total size of all jobs not marked as a final result or ancestor of final result

### `get_num_active_licenses()`

Get number of acquired licenses for running jobs

* **Returns:** number of active licenses
* **Return type:** int

### `get_project(project_uid: str, \*args)`

Get information about a single project

* **Parameters:**
  * **project\_uid** (*str*) -- the id of the project
  * **args** -- extra args (comma seperated) that contain the keys (if any) to project when returning the mongodb doc
* **Returns:** the information related to the project thats stored in the database
* **Return type:** str

### `get_project_dir_abs(project_uid: str)`

Get the project's absolute directory with all environment variables in the path resolved

* **Parameters:** **project\_uid** (*str*) -- target project UID, e.g., "P3"
* **Raises:** **ValueError** -- when the project does not exist
* **Returns:** the absolute path to the project directory
* **Return type:** str

### `get_project_jobs_by_status(project_uid, workspace_uid=None, statuses=[], fields=[])`

Get all jobs matching the given statuses, optionally scoped to a workspace

### `get_project_symlinks(project_uid: str)`

Get all symbolic links in the given project directory

* **Parameters:** **project\_uid** (*str*) -- target project UID, e.g., "P3"
* **Returns:** List symlink paths, either absolute or relative to project directory
* **Return type:** str

### `get_project_title_slug(project_title: str)`

Returns a slugified version of a project title

* **Parameters:** **project\_title** (*str*) -- Requested project title
* **Returns:** URL- and file-system- safe slug of the given project title
* **Return type:** str

### `get_result_download_abs_path(project_uid: str, result_spec: str)`

Get the absolute path to the dataset for the given result type

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **result\_spec** (*str*) -- result slot name e.g., "J3.particles.blob"
* **Returns:** absolute path
* **Return type:** str

### `get_running_version()`

Get the current CryoSPARC version

* **Returns:** e.g., v3.3.2
* **Return type:** str

### `get_runtime_diagnostics()`

Get runtime diagnostics for the CryoSPARC instance

* **Returns:** formatted with diagnostics
* **Return type:** dict

### `get_scheduler_lanes()`

Returns a list of lanes that are registered with the master scheduler

* **Returns:** list of information about each lane
* **Return type:** dicts

### `get_scheduler_target_cluster_info(name: str)`

Get cluster information for given cluster name

* **Parameters:** **name** (*str*) -- name of cluster/lane (lane must be of type 'cluster')
* **Returns:** json string containing cluster information
* **Return type:** str
* **Raises:** AssertionError

### `get_scheduler_target_cluster_script(name: str)`

Get cluster job submission template for specified cluster

* **Parameters:** **name** (*str*) -- name of cluster/lane (lane must be of type 'cluster')
* **Returns:** job scheduler script template
* **Return type:** str
* **Raises:** AssertionError

### `get_scheduler_targets()`

Returns a list of worker nodes that are registered with the master scheduler

* **Returns:** list of information about each worker node
* **Return type:** dicts

### `get_supervisor_log_latest(service: str, n: int = 50)`

Read the last n lines from the given service log. Service must be one of the following:

* `"app"`
* `"database"`
* `"command_core"`
* `"command_vis"`
* `"command_rtp"`
* `"app_api"`
* `"app_legacy"`
* `"supervisord"`
* **Parameters:**
  * **service** (*str*) -- Service to retrive logs for
  * **n** (*int*\*,\* *optional*) -- Number of lines to retrieve, defaults to 50
* **Returns:** logs as a string separated by new lines
* **Return type:** str

### `get_system_info()`

Returns system-related information related to the cryosparc app

* **Returns:** dictionary listing information about cryosparc environment
* **Return type:** dict

### `get_tag_count_by_type()`

Get a dictionary of where keys are tag types and values are how many tags there are of each type

* **Returns:** dict of integers
* **Return type:** dict

### `get_tag_counts(tag_uid: str)`

Get a dictionary of counts contining how many entities (e.g., projects, jobs) are tagged with the given tag UID

* **Parameters:** **tag\_uid** (*str*) -- target tag UID, e.g., "T1"
* **Returns:** counts organized by entity type
* **Return type:** str

### `get_tags_by_type()`

Get all tags as a dictionary organized by tag type. Each key is a type such as "general" or "job" and each value is a list of tags

* **Returns:** dict of lists of tags
* **Return type:** dict

### `get_tags_of_type(tag_type: str)`

Get a list of tags with the given type

* **Parameters:** **tag\_type** (*str*) -- tag type such as "general", "project", "job", etc.
* **Returns:** List of tags with the given type
* **Return type:** list

### `get_targets_by_lane(lane_name, node_only=False)`

Returns a list of worker nodes that are registered with the master scheduler and are in the given lane

### `get_user_default_priority(email_address: str)`

Get a user's priority when launching jobs. Defaults to the instance's configured job priority if not set.

* **Parameters:** **email\_address** (*str*) -- email for the target user
* **Returns:** the job priority, defaulting to 0 if not set
* **Return type:** int

### `get_user_lanes(email_address: str)`

Retrieve a list of lens the user with the given email address is allowed to queue jobs to.

* **Parameters:** **email\_address** (*str*) -- target user account email address
* **Returns:** list of lanes
* **Return type:** list

### `get_user_tags(user_id: str)`

Get a list of tags created by the given user account ID

* **Parameters:** **user\_id** (*str*) -- target user account ID
* **Returns:** list of tags
* **Return type:** list

### `get_worker_nodes()`

Returns a list of worker nodes registered in the master scheduler

* **Returns:** list of information about each worker node
* **Return type:** dicts

### `get_workspace(project_uid, workspace_uid, \*args)`

A helper function that returns the mongodb document of the workspace requested, only including the fields passed

* **Parameters:**
  * **project\_uid** (*str*) -- the uid of the project that contains the workspace
  * **workspace\_uid** (*str*) -- the uid of the workspace to retrieve
  * **args** -- extra args (comma seperated) that contain the keys (if any) to project when returning the mongodb doc
* **Returns:** dict - the mongodb workspace document

### `import_job(owner_user_id: str, project_uid: str, workspace_uid: str, abs_path_export_job_dir: str)`

Import the given exported job directory into the given project. Exported job directory must exist in the project directory. By convention, may be added into the "imports" directory

* **Parameters:**
  * **owner\_user\_id** (*str*) -- \_id of user performing this import operation
  * **project\_uid** (*str*) -- project UID to import into, e.g., "P3"
  * **workspace\_uid** (*str*) -- workspace UID to import into, e.g., "W1"
  * **abs\_path\_export\_job\_dir** (*str*) -- path to exported job directory. Must be inside the project directory.

### `import_jobs(jobs_manifest, abs_path_export_project_dir, new_project_uid, owner_user_id, notification_id)`

Imports jobs using the job manifest

* **Parameters:**
  * **jobs\_manifest** (*str*) -- the job data loaded from a job\_manifest.json file
  * **abs\_path\_export\_project\_dir** (*str*) -- the import project directory
  * **new\_project\_uid** (*str*) -- uid of the project to import the jobs into
  * **owner\_user\_id** (*str*) -- the id of the user importing the jobs
  * **notification\_id** (*str*) -- the import project notification uid that will have its progress meter updated as import completes

### `import_project(owner_user_id: str, abs_path_export_project_dir: str)`

Import any project directory that was previously exported by CryoSPARC

* **Parameters:**
  * **owner\_user\_id** (*str*) -- the mongo object id ("\_id") of the user requesting to import the project
  * **abs\_path\_export\_project\_dir** (*str*) -- the absolute project directory containing the jobs to import

### `import_workspaces(workspaces_doc_data, abs_path_export_project_dir, new_project_uid, owner_user_id, notification_id)`

Imports workspaces and live sessions

* **Parameters:**
  * **workspaces\_doc\_data** (*str*) -- workspace data loaded from workspaces.json file
  * **abs\_path\_export\_project\_dir** (*str*) -- the import project directory
  * **new\_project\_uid** (*str*) -- uid of the project to import the workspaces into
  * **owner\_user\_id** (*str*) -- the id of the user importing the workspaces
  * **notification\_id** (*str*) -- import project notification uid that will have its progress meter updated as import completes

### `is_admin(user_id: str)`

Returns True if the given user account ID has admin privileges

* **Parameters:** **user\_id** (*str*) -- \_id property of the target user
* **Returns:** Whether the user is an admin
* **Return type:** bool

### `job_add_to_workspace(project_uid: str, job_uid: str, workspace_uid: str)`

Adds a job to the specified workspace

* **Parameters:**
  * **project\_uid** (*str*) -- the id of the project
  * **job\_uid** (*str*) -- the id of the job to add
  * **workspace\_uid** (*str*) -- the id of the workspace
* **Returns:** the number of modified documents
* **Return type:** int

### `job_cart_create(project_uid: str, workspace_uid: str, created_by_user_id: str, output_result_groups: list, new_job_type: str)`

Given a list of output result groups and a job type, create the new job, and connect the output result groups to the input slots of the newly created job.

* **Parameters:**
  * **project\_uid** (*str*) -- project ID, e.g., P3
  * **workspace\_uid** (*str*) -- workspace ID, e.g., W1
  * **created\_by\_user\_id** (*str*) -- \_id of user
  * **output\_result\_groups** (*list*) -- output result groups to connect to the input groups of the new job
  * **new\_job\_type** (*str*) -- the job type to create

### `job_clear_param(project_uid: str, job_uid: str, param_name: str)`

Reset the given parameter to its default value.

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **job\_uid** (*str*) -- target job UID, e.g.,
  * **param\_name** (*str*) -- target parameter name, e.g., "refine\_symmetry"
* **Returns:** whether the job has an build errors
* **Return type:** bool

### `job_clear_streamlog(project_uid: str, job_uid: str)`

Delete all entries from the given job's event log

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **job\_uid** (*str*) -- target job UID, e.g., "J42"
* **Returns:** delete result, including affected document count
* **Return type:** dict

### `job_connect_group(project_uid: str, source_group: str, dest_group: str)`

Connect the given source output group to the target input group. Each group must be formatted as `<job uid>.<group name>`

* **Parameters:**
  * **project\_uid** (*str*) -- currenct project UID, e.g., "P3"
  * **source\_group** (*srr*) -- source output group, e.g., "J1.movies"
  * **dest\_group** (*str*) -- destination input group, e.g., "J2.exposures"
* **Returns:** whether the job has any errors
* **Return type:** bool

### `job_connect_result(project_uid: str, source_result: str, dest_slot: str)`

Connect the given source output slot to the target input slot, where the main group (e.g., `particles`) remains but a set of fields (e.g., `particles.blob`) is added or replaced.

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **source\_result** (*str*) -- output slot to connect, e.g., "J1.particles.blob"
  * **dest\_slot** (*str*) -- input slot to connect to, e.g., "J2.particles.0.blob"
* **Returns:** whether the job has any build errors
* **Return type:** bool

### `job_connected_group_clear(project_uid: str, dest_group: str, connect_idx: int)`

Clear the given job input group. Group must be formatted as `<job uid>.<group name>`

* **Parameters:**
  * **project\_uid** (*str*) -- current project UID, e.g., "P3"
  * **dest\_group** (*str*) -- Group to clear, e.g., "J2.exposures"
  * **connect\_idx** (*int*) -- Connection index to clear when multiple outputs are connected to the same input. Set to 0 to clear the first input.
* **Returns:** whether the job has any build errors
* **Return type:** bool

### `job_connected_result_clear(project_uid: str, dest_slot: str, passthrough=False)`

Clear a given input slot from a job, where the main group (e.g., `particles`) is left intact but a set of fields (e.g., `particles.blob`) is removed

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **dest\_slot** (*str*) -- input slot to clear, e.g., "J2.particles.0.blob"
  * **passthrough** (*bool*\*,\* *optional*) -- set to True the input slot was passed through from its parent output (though not strictly necessary), defaults to False
* **Returns:** whether the job has any build errors
* **Return type:** bool

### `job_find_ancestors(project_uid: str, job_uid: str)`

Find all jobs that provided inputs to the given job by following the given job's inputs backward up the job tree.

* **Parameters:**
  * **project\_uid** (*str*) -- project UID where jobs reside, e.g., "P3"
  * **job\_uid** (*str*) -- uid of job, e.g., "J42"
* **Returns:** sorted list of job UIDs
* **Return type:** list

### `job_find_descendants(project_uid: str, job_uid: str)`

Find all jobs that the given job provided outputs to by following the given job's outputs forward down the job tree.

* **Parameters:**
  * **project\_uid** (*str*) -- project UID where jobs reside, e.g., "P3"
  * **job\_uid** (*str*) -- uid of job, e.g., "J42"
* **Returns:** sorted list of job UIDs
* **Return type:** list

### `job_has_build_errors(project_uid: str, job_uid: str)`

Whether the given job has any build errors

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **job\_uid** (*str*) -- target job UID, e.g., "J42"
* **Returns:** True or False
* **Return type:** bool

### `job_import_replace_symlinks(project_uid: str, job_uid: str, prefix_cut: str, prefix_new: str)`

Update symbolic links to imported data in the given job directory when the original data has been moved.

* **Parameters:**
  * **project\_uid** (*str*) -- target project uid, e.g., "P3"
  * **job\_uid** (*str*) -- uid of target job, e.g. "J42"
  * **prefix\_cut** (*str*) -- Old path prefix of external data, e.g., "/path/to/dataold"
  * **prefix\_new** (*str*) -- New path prefix where external data has been moved, e.g., "/path/to/datanew"
* **Returns:** number of replaced symbolic links
* **Return type:** str

### `job_rebuild(project_uid: str, job_uid: str)`

Re-run the builder for the given job to re-generate parameters and input slots

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID
  * **job\_uid** (*str*) -- target job UID

### `job_remove_from_workspace(project_uid: str, job_uid: str, workspace_uid: str)`

Removes a job from the specified workspace

* **Parameters:**
  * **project\_uid** (*str*) -- the id of the project
  * **job\_uid** (*str*) -- the id of the job to remove
  * **workspace\_uid** (*str*) -- the id of the workspace
* **Returns:** the number of modified documents
* **Return type:** int

### `job_send_streamlog(project_uid: str, job_uid: str, message: str, error: bool = False, flags: List[str] = [], imgfiles: List[EventLogAsset] = [])`

Add the given message to the target job's event log

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **job\_uid** (*str*) -- target job uid, e.g., "J42"
  * **message** (*str*) -- Message to log
  * **error** (*bool*\*,\* *optional*) -- Whether to show as error, defaults to False
  * **flags** (*list*\*\[**str**]\*\*,\* *optional*) -- Additional event flags
  * **imgfiles** (*list*\*\[**StreamlogAsset**]\*) -- Uploaded GridFS files to attach to event
* **Returns:** Created mongo event ID
* **Return type:** str

### `job_set_param(project_uid: str, job_uid: str, param_name: str, param_new_value: Any)`

Set the given job parameter to the given value

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **job\_uid** (*str*) -- target job UID, e.g., "J42"
  * **param\_name** (*str*) -- target parameter name, e.g., "random\_seed"
  * **param\_new\_value** (*any*) -- new parameter value
* **Returns:** whether the job has any build errors
* **Return type:** bool

### `kill_job(project_uid, job_uid, killed_by_user_id=None)`

Kill the given running job

* **Parameters:**
  * **project\_uid** (*str*) -- uid of the project that contain the job to kill
  * **job\_uid** (*str*) -- job uid to kill
  * **killed\_by\_user\_id** (*str*) -- ID of user that killed the job, optional
* **Raises:** AssertionError

### `list_projects()`

Get list of all projects available

* **Returns:** all projects available in the database
* **Return type:** list

### `list_users()`

Show a table of all CryoSPARC user accounts

### `list_workspaces(project_uid=None)`

List all workspaces inside a given project (or all projects if not specified)

* **Parameters:** **project\_uid** (*str*\*,\* *optional*) -- target project UID, e.g. "P1", defaults to None
* **Returns:** list of workpaces in project or all projects if not specified
* **Return type:** list

### `make_job(job_type: str, project_uid: str, workspace_uid: str, user_id: str, created_by_job_uid: str | None = None, title: str | None = None, desc: str | None = None, params: dict = {}, input_group_connects: dict = {}, enable_bench: bool = False, priority: int | None = None, do_layout: bool = True)`

Create a new job with the given type in the given project/workspace

To see all available job types, see cryosparc\_compute/jobs/register.py (look for the value of the `contains` keys).

To see what parameters are available for a job and what values are available, reference the build.py file in cryosparc\_compute/jobs that pertains to the desired job type.

* **Parameters:**
  * **job\_type** (*str*) -- Type of job
  * **project\_uid** (*str*) -- project ID, e.g., P3
  * **workspace\_uid** (*str*) -- workspace ID, e.g., W1
  * **user\_id** (*str*) -- ID of user
  * **created\_by\_job\_uid** (*str*\*,\* *optional*) -- ID of the parent job that created this, defaults to None
  * **title** (*str*\*,\* *optional*) -- Descriptive title for what this job is for, defaults to None
  * **desc** (*str*\*,\* *optional*) -- Detailed description, defaults to None
  * **params** (*dict*\*,\* *optional*) -- Parameter settings for the job, defaults to {}
  * **input\_group\_connects** (*dict*\*,\* *optional*) -- Connected input groups, which each key is name and each value has format `JXX.output_group_name`, defaults to {}
  * **enable\_bench** (*bool*\*,\* *optional*) -- enable benchmarking stats for this job, defaults to False
  * **priority** (*int*\*,\* *optional*) -- job priority, defaults to None (use default priority)
  * **do\_layout** (*bool*\*,\* *optional*) -- re-compute the workspace and project tree view after creating this job, defaults to True
* **Returns:** ID of new job
* **Return type:** str
* **Example:**

The following makes a 3-class ab-initio job with particles connected from a Select 2D Classes job:

```python
juid = cli.make_job(
    job_type='homo_abinit', project_uid='P3', workspace_uid='W1',
    user_id=cli.GetUser('email@example.com')['_id'], params={'abinit_K':
    3}, input_group_connects={'particles': 'J41.particles_selected'})

print(juid)  # "J42"
```

### `parse_template_vars(target)`

Parses and returns a list of template variable names from the submission script and cluster commands in a cluster target :param target: a cluster target :return: list of template variable names :rtype: list\[str]

### `propose_clone_job_chain(project_uid: str, start_job_uid: str, end_job_uid: str)`

Deprecated - Old function name for get\_job\_chain() which was used only for cloning jobs

### `refresh_job_types()`

Reload the available input and parameter types on jobs created prior to a CryoSPARC update. Run this following a CryoSPARC update to ensure jobs run with the correct parameters.

* **Returns:** list of the results of the mongodb query
* **Return type:** dicts

### `remove_scheduler_lane(name: str)`

Removes the specified lane and any targets assigned under the lane in the master scheduler

#### `NOTE`

This will remove any worker node associated with the specified lane.

* **Parameters:** **name** (*str*) -- the name of the lane to remove

### `remove_scheduler_target_cluster(name: str)`

Removes the specified cluster/lane and any targets assigned under the lane in the master scheduler

#### `NOTE`

This will remove any worker node associated with the specified cluster/lane.

* **Parameters:** **name** (*str*) -- the name of the cluster/lane to remove
* **Returns:** "True" if successful
* **Return type:** bool

### `remove_scheduler_target_node(hostname: str)`

Removes a target worker node from the master scheduler

* **Parameters:** **hostname** (*str*) -- the hostname of the target worker node to remove

### `remove_tag(tag_uid: str)`

Delete the given tag and remove it from all connected entitites

* **Parameters:** **tag\_uid** (*str*) -- target tag UID, e.g., "T1"
* **Returns:** contains deleted count
* **Return type:** dict

### `remove_tag_from_job(project_uid: str, job_uid: str, tag_uid: str)`

Remove the given tag from the job

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **job\_uid** (*str*) -- target job UID, e.g., "W1"
  * **tag\_uid** (*str*) -- target tag UID, e.g., "T1"
* **Returns:** contains modified jobs count
* **Return type:** dict

### `remove_tag_from_project(project_uid: str, tag_uid: str)`

Remove the given tag from the given projet

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **tag\_uid** (*str*) -- target tag UID, e.g., "T1"
* **Returns:** contains modified project count
* **Return type:** dict

### `remove_tag_from_session(project_uid: str, session_uid: str, tag_uid: str)`

Remove the given tag from the session

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **session\_uid** (*str*) -- target session UID, e.g., "W1"
  * **tag\_uid** (*str*) -- target tag UID, e.g., "T1"
* **Returns:** contains modified sessions count
* **Return type:** dict

### `remove_tag_from_workspace(project_uid: str, workspace_uid: str, tag_uid: str)`

Remove the given tag from the workspace

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **workspace\_uid** (*str*) -- target workspace UID, e.g., "W1"
  * **tag\_uid** (*str*) -- target tag UID, e.g., "T1"
* **Returns:** contains modified workspaces count
* **Return type:** dict

### `request_delete_project(project_uid: str, request_user_id: str)`

Confirm what jobs and workspaces will be deleted when requesting to delete a project. Does not perform the deletion (see `delete_project`).

* **Parameters:**
  * **project\_uid** (*str*) -- uid of the project to be deleted
  * **request\_user\_id** (*str*) -- \_id of user requesting the project to be deleted
* **Returns:** If the project does not have any jobs or workspaces related to it, will "disable" the project and return a string confirmation.
* **Return type:** str
* **Returns:** If the project has jobs or workspaces associated with it, returns 2 lists. First list are all non-deleted jobs that exist in the project, and second list are all workspaces within the project to be deleted.
* **Return type:** tuple

### `request_delete_workspace(project_uid: str, workspace_uid: str, request_user_id: str)`

Confirm what jobs be deleted when requesting to delete a workspace. Call `delete_workspace` with results to perform deletion

* **Parameters:**
  * **project\_uid** (*str*) -- uid of the project containing the workspace to be deleted
  * **workspace\_uid** (*str*) -- uid of the workspace to be deleted
* **Returns:** If the workspace does not have any jobs related to it, "disables" the workspace and returns a string confirmation.
* **Return type:** str
* **Returns:** If the workspace has jobs associated with it, returns 2 lists. First list includes jobs only within the given workspace and can be deleted, and second list includes jobs that are a part of other workspaces and should not be deleted.
* **Return type:** tuple

### `request_reset_password(email)`

Generate a password reset token for a user with the given email. The token will appear in the Admin > User Management interface.

* **Parameters:** **email** (*str*) -- email address of target user account

### `reset_password(email, newpass)`

Reset a cryosparc user's password

* **Parameters:**
  * **email** (*str*) -- the user's email address
  * **password** (*str*) -- the user's new password
* **Returns:** the number of modified documents
* **Return type:** str

### `run_external_job(project_uid: str, job_uid: str, status: typing_extensions.Literal[running, waiting] = 'waiting')`

Special run method which marks an External job as running or waiting.

* **Parameters:**
  * **project\_uid** (*str*) -- Project UID of target job
  * **job\_uid** (*str*) -- Job UI
  * **status** (*str*\*,\* *optional*) -- Status to run with, defaults to "waiting"

### `save_extensive_validation_benchmark_data(instance_information: dict, job_timings: dict)`

Save extensive validation benchmark data to database for visulisation in UI

### `save_performance_benchmark_references()`

On every startup, this function is called to repopulate the benchmark\_references collection.

### `set_cluster_job_custom_vars(project_uid, job_uid, cluster_job_custom_vars)`

Set a job's custom variables for cluster submission

### `set_instance_banner(active: bool, title: str | None = None, body: str | None = None)`

Set an instance banner as active or inactive. Updates title and body if provided.

* **Parameters:**
  * **active** (*bool*) -- True/False to activate/deactivate
  * **title** (*str*\*,\* *optional*) -- Banner title, defaults to None
  * **body** (*str*\*,\* *optional*) -- Banner body, defaults to None
* **Returns:** title and body
* **Return type:** dict

### `set_instance_default_job_priority(priority)`

Set the default priority for jobs queued by users without an explicit priority set in Admin > User Management

* **Parameters:** **priority** (*int*) -- Non-negative priority number

### `set_job_final_result(project_uid: str, job_uid: str, is_final_result: bool)`

Sets job final result flag and updates flags for all jobs in the project

### `set_login_message(active: bool, title: str | None = None, body: str | None = None)`

Set an login message as active or inactive. Updates title and body if provided.

* **Parameters:**
  * **active** (*bool*) -- True/False to activate/deactivate
  * **title** (*str*\*,\* *optional*) -- Login message title, defaults to None
  * **body** (*str*\*,\* *optional*) -- Login message body, defaults to None
* **Returns:** title and body
* **Return type:** dict

### `set_maintenance_mode(maintenance_mode: bool)`

Set maintenance mode status.

* **Parameters:** **maintenance\_mode** (*bool*) -- True to enable, False to disable

### `set_project_owner(project_uid: str, user_id: str)`

Updates a project's owner

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID to update, e.g., "P3"
  * **user\_id** (*str*) -- the owner's user id

### `set_project_param_default(project_uid: str, param_name: str, value: Any)`

Set a default value for a given parameter name globally for the given project

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **param\_name** (*str*) -- target parameter name, e.g., "compute\_use\_ssd"
  * **value** (*Any*) -- target parameter default value

### `set_scheduler_target_node_cache(hostname: str, cache_reserve: int | None = None, cache_quota: int | None = None)`

Sets the cache reserve and cache quota for a target worker node

* **Parameters:**
  * **hostname** (*str*) -- the hostname of the target worker node
  * **cache\_reserve** (*int*\*,\* *optional*) -- the size (in MB) to reserve on the SSD for the cryosparc cache
  * **cache\_quota** (*int*\*,\* *optional*) -- the max size (in MB) to use on the SSD for the cryosparc cache
* **Raises:** AssertionError

### `set_scheduler_target_node_lane(hostname, lane)`

Sets the lane of a target worker node

* **Parameters:**
  * **hostname** (*str*) -- the hostname of the target worker node
  * **lane** (*str*) -- the name of the lane to assign to the target worker node:
* **Returns:** target worker node's updated configurations
* **Return type:** list
* **Raises:** AssertionError

### `set_scheduler_target_property(hostname: str, key: str, value: Any)`

Set a property for the target worker node

* **Parameters:**
  * **hostname** (*str*) -- the hostname of the target worker node
  * **key** (*str*) -- the key of the property whose value is being modified
  * **value** (*str*) -- the actual value to set for the property
* **Returns:** information about all targets
* **Return type:** list
* **Raises:** AssertionError

### `set_user_allowed_prefix_dir(user_id: str, allowed_prefix: str)`

Sets directories that users are allowed to query from the file browser

* **Parameters:**
  * **user\_id** (*str*) -- the mongo id of the user to update
  * **allowed\_prefix** (*str*) -- the path of the directory the user can query inside (must start with "/", and must be an absolute path)
* **Returns:** True if successful
* **Return type:** bool
* **Raises:** AssertionError

### `set_user_default_priority(email_address: str, priority: int)`

Set a user's priority when launching jobs. This is equivalent to changing the Default Job Priority in the Admin > User Management interface.

* **Parameters:**
  * **email\_address** (*\_type\_*) -- \_description\_
  * **priority** (*\_type\_*) -- \_description\_

### `set_user_lanes(email_address: str, assigned_lanes: List[str])`

Only allow a user account with the given email address to queue to the given lanes.

* **Parameters:**
  * **email\_address** (*str*) -- target user account email
  * **assigned\_lanes** (*list*) -- list of lane names to assign

### `set_user_state_var(user_id: str, key: str, value: Any, set_on_insert_only: bool = False)`

Updates a user's state variable

* **Parameters:**
  * **user\_id** (*str*) -- the user's id
  * **key** (*str*) -- the name of the key to update or insert
  * **value** (*str*) -- the actual value to insert under the key in the state variable
  * **set\_on\_insert\_only** (*bool*\*,\* *optional*) -- specifies if `setOnInsert` should be used (if the update op results in an insertion of a document, this will assign the specified values to the fields in the doc), defaults to False
* **Returns:** confirmation that the update happened successfully
* **Return type:** bool

### `start_worker_test(project_uid: str, test_type: str = 'all', targets: list | None = None, verbose: bool = True)`

Launch a test to ensure CryoSPARC can correctly queue to all or only the given list of targets

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID to launch test int
  * **test\_type** (*str*\*,\* *optional*) -- 'all', 'launch', 'ssd' or 'gpu', defaults to 'all'
  * **targets** (*list*\*,\* *optional*) -- list of targets to test, omit to test all, defaults to None
  * **verbose** (*bool*\*,\* *optional*) -- if True, show extended log information, defaults to True

### `take_over_project(project_uid, force=False)`

Write a lockfile to an existing project so that no CryoSPARC instances outside of the current one may modify it.

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **force** (*bool*\*,\* *optional*) -- If True, (over-)write the lock file even if a lock file already exists, defaults to False
* **Raises:** **Exception** -- If lock file cannot be written
* **Returns:** True if takeover succeeded
* **Return type:** bool

### `take_over_projects(force=False)`

Write lockfiles to all existing projects so that no CryoSPARC instances outside of the current one may modify it.

* **Parameters:** **force** (*bool*\*,\* *optional*) -- Takeover all projects even if they already have a lock file from a diffeent instance, defaults to False

### `test_authentication(project_uid, job_uid)`

Test if a worker running a job can correctly authenticate with command\_core

### `test_connection(sleep_time=0)`

Check the connection to the command\_core service. Returns True if the connection succeeded.

* **Parameters:** **sleep\_time** (*float*\*,\* *optional*) -- How long to sleep for, in seconds, defaults to 0
* **Returns:** True if connection succeeded
* **Return type:** bool

### `unarchive_project(project_uid: str, abs_path_to_project_dir: str)`

Reverse archive operation.

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **abs\_path\_to\_project\_dir** (*str*) -- Project directory to unarchive from

### `unset_project_param_default(project_uid: str, param_name: str)`

Clear the per-project default value for the given parameter name.

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **param\_name** (*str*) -- target parameter name, e.g., "compute\_use\_ssd"

### `update_all_job_sizes(asynchronous=True, sleep_time=0)`

Recompute the folder sizes of all jobs

* **Parameters:**
  * **asynchronous** (*bool*\*,\* *optional*) -- if True, returns immediately and computes sizes in the background, defaults to True
  * **sleep\_time** (*float*\*,\* *optional*) -- if asynchronous is True, waits the given number of seconds before proceeding, defaults to 0

### `update_final_result_statuses(project_uid: str)`

Update jobs in a project to have the correct flags set when marked as final or as an ancestor of marked as final

### `update_job(project_uid: str, job_uid: str, attrs: dict, operation='$set')`

Update a specific job's document with the provided attributes

* **Parameters:**
  * **project\_uid** (*str*) -- project UID
  * **job\_uid** (*str*) -- uid of job
  * **attrs** (*dict*) -- the attribute to modify in the job document
  * **operation** (*str*\*,\* *optional*) -- the operation to perform on the document, defaults to '$set'

### `update_parents_and_children_for_project(project_uid)`

Restore tree view if broken following an import

* **Parameters:** **project\_uid** (*str*) -- target project UID, e.g., "P3"

### `update_project(project_uid: str, attrs: dict, operation='$set', export=True)`

Update a project's attributes

* **Parameters:**
  * **project\_uid** (*str*) -- the id of the project to update
  * **attrs** (*dict*) -- the key to update and its value
  * **operation** (*str*) -- the type of mongodb operation to perform
* **Example:**

```python
update_project(project_uid, {'project_dir' : new_abs_path})
```

### `update_project_directory(project_uid: str, new_project_dir: str)`

Safely updates the project directory of a project given a directory. Checks if the directory exists, is readable, and writeable.

* **Parameters:**
  * **project\_uid** (*str*) -- uid of the project to update
  * **new\_project\_dir\_container** (*str*) -- the new directory

### `update_project_root_dir(project_uid: str, new_project_dir_container: str)`

Updates the root directory of a project (creates a new project (PXXX) folder within the new root directory and updates the project document)

#### `NOTE`

the root project directory passed cannot contain a folder with the same project (PXXX) folder name

* **Parameters:**
  * **project\_uid** (*str*) -- uid of the project to update
  * **new\_project\_dir\_container** (*str*) -- the new "root" path where the project (PXXX) folder is to be created

### `update_project_size(project_uid: str, use_prt: bool = True)`

Calculates the size of the project. Similar to running du -sL inside the project dir

* **Parameters:** **project\_uid** (*str*) -- Unique ID of project to update, e.g., "P3"

### `update_session(project_uid: str, session_uid: str, attrs: dict, operation='$set')`

Similar to `update_workspace`, but for sessions

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **session\_uid** (*str*) -- target session UID, e.g., "S5"
  * **attrs** (*dict*) -- Attributes to update
  * **operation** (*str*\*,\* *optional*) -- mongo operation type, defaults to '$set'

### `update_tag(tag_uid: str, title: str | None = None, colour: str | None = None, description: str | None = None)`

Update the title, colour and/or description of the given tag UID

* **Parameters:**
  * **tag\_uid** (*str*) -- target tag UID, e.g., "T1"
  * **title** (*str*\*,\* *optional*) -- new value for title, defaults to None
  * **colour** (*str*\*,\* *optional*) -- new value of colour, defaults to None
  * **description** (*str*\*,\* *optional*) -- new value of description, defaults to None
* **Returns:** updated tag document
* **Return type:** dict

### `update_user(email: str, password: str, username: str | None = None, first_name: str | None = None, last_name: str | None = None, admin: bool | None = None)`

Updates a cryosparc user's details. Email and password are required, other params will only be set if they are not empty.

* **Parameters:**
  * **email** (*str*) -- the user's email address
  * **password** (*str*) -- the user's password
  * **username** (*str*\*,\* *optional*) -- new username of the user
  * **first\_name** (*str*\*,\* *optional*) -- new given name of the user
  * **last\_name** (*str*\*,\* *optional*) -- new surname of the user
  * **admin** (*bool*\*,\* *optional*) -- whether the user should be admin or not
* **Returns:** confirmation message if successful or not
* **Return type:** str

### `update_workspace(project_uid: str, workspace_uid: str, attrs: dict, operation='$set', export=True)`

Update properties for the given workspace

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **workspace\_uid** (*str*) -- target workspace UID, e.g., "W4"
  * **attrs** (*dict*) -- Attributes to update
  * **operation** (*str*\*,\* *optional*) -- mongo operation type, defaults to '$set'
  * **export** (*bool*\*,\* *optional*) -- Whether to dump this workspace to disk export to other instances, defaults to True

### `upload_extensive_validation_benchmark_data(instance_information: dict, job_timings: dict)`

Upload extensive validation benchmark data

### `validate_enqueue_job(project_uid, job_uid, interactive, lightweight, lane, hostname)`

Validate a job's queueing configuration

### `validate_license()`

Check whether the active license ID is valid. If the license is invalid, use `dump_license_validation_results()` to get additional details.

* **Returns:** True if valid, False otherwise.
* **Return type:** bool

### `validate_project_creation(project_title: str, project_container_dir: str)`

Validate project title and container directory for a project creation request.

* **Parameters:**
  * **project\_title** (*str*) -- title of the project
  * **project\_container\_dir** (*str*) -- container directory of the project
* **Returns:** dict with a slug of the project name, validity, and a message if the request was invalid
* **Return type:** dict

### `verify_cluster(name: str)`

Ensure cluster has been properly configured by executing a generic 'info' command

* **Parameters:** **name** (*str*) -- name of cluster/lane (lane must be of type 'cluster')
* **Returns:** the result of the command execution
* **Return type:** str
* **Raises:** AssertionError

### `wait_job_complete(project_uid: str, job_uid: str, timeout: int | float = 5)`

Hold the request and prevent from returning until the given job status is "completed", the given timeout is reached or the global command client request timeout is reached (default 300 seconds)

* **Parameters:**
  * **project\_uid** (*str*) -- target project UID, e.g., "P3"
  * **job\_uid** (*str*) -- target job UID, e.g., "J42"
  * **timeout** (*float*\*,\* *optional*) -- how long to wait in seconds, defaults to 5
* **Returns:** the job status when the timeout is reached
* **Return type:** str

### `get_email_by_id(user_id: str)`

Get the registered email address for the user account with the given ID

* **Parameters:** **user\_id** (*str*) -- \_id property of the target user account
* **Returns:** first email property in the database, or safe default value if not found
* **Return type:** str

### `get_id_by_email(email: str)`

Get the \_id property for the user account with the given email address

* **Parameters:** **email** (*str*) -- target user email address
* **Raises:** **ValueError** -- If user is not found
* **Returns:** \_id property of the target user
* **Return type:** str

### `get_id_by_email_password(email: str, password: str)`

Retrieve the ID for a user given their email and password. Raises error if matching email/password combo is not found.

* **Parameters:**
  * **email** (*str*) -- Target user email address
  * **password** (*str*) -- Target user password
* **Returns:** ID of requesting user
* **Return type:** str

### `get_user_id(user: str)`

Get a user's ID by one of its unique identifiers, including email, username and the ID itself. Returns None if user ID is not found or given identifier is not unique.

### `get_username_by_id(user_id: str)`

Get the registered user name for the user account with the given ID

* **Parameters:** **user\_id** (*str*) -- \_id property of the target user account
* **Returns:** name property in the database
* **Return type:** str


# cryosparcw reference (≤v4.7)

How to use the cryosparcw utility for managing CryoSPARC workers

{% hint style="warning" %}
This page refers to CryoSPARC ≤v4.7.

For v5.0+, please see [cryosparcw reference (v5.0+)](/setup-configuration-and-management/management-and-monitoring-v5.0/cryosparcw-reference-v5.0)
{% endhint %}

## Worker Management with `cryosparcw`

Run all commands in this section while logged into the workstation or worker nodes where the `cryosparc_worker` package is installed.

Verify that `cryosparcw` is in the worker's `PATH` with this command:

```
which cryosparcw
```

and ensure the output is not empty.

Alternatively navigate to the worker installation directory and in `bin/cryosparcw`. Example:

```bash
cd /path/to/cryosparc_worker
bin/cryosparcw gpulist
```

### `cryosparcw env`

Equivalent to [`cryosparcm env`](#cryosparcm-env), but for worker nodes.

### `cryosparcw call <command>`

Execute a shell command in a *transient* CryoSPARC worker shell environment. For example,

`cryosparcw call which python`

prints the path of the CryoSPARC worker environment’s Python executable.

### `cryosparcw connect <options>`

Run this command on the worker node that you wish to register with the master, or whose existing registration you wish to update.

Enter `cryosparcw connect --help` to see full usage details.

```bash
$ cryosparcw connect --help
usage: connect.py [-h] [--worker WORKER] [--master MASTER] [--port PORT]
                  [--sshstr SSHSTR] [--update] [--nogpu] [--gpus GPUS]
                  [--nossd] [--ssdpath SSDPATH] [--ssdquota SSDQUOTA]
                  [--ssdreserve SSDRESERVE] [--lane LANE] [--newlane]
                  [--rams RAMS] [--cpus CPUS] [--monitor_port MONITOR_PORT]

Connect to cryoSPARC master

optional arguments:
  -h, --help            show this help message and exit
  --worker WORKER
  --master MASTER
  --port PORT
  --sshstr SSHSTR
  --update
  --nogpu
  --gpus GPUS
  --nossd
  --ssdpath SSDPATH
  --ssdquota SSDQUOTA
  --ssdreserve SSDRESERVE
  --lane LANE
  --newlane
  --rams NUM_RAMS
  --cpus NUM_CPUS
  --monitor_port MONITOR_PORT
```

Example command to connect a worker on a new resources lane.

```bash
cryosparcw connect \
    --worker $(hostname -f) \
    --master csmaster.local \
    --port 61000 \
    --ssdpath /scratch/cryosparc_cache \
    --lane $(hostname -s) \
    --newlane
```

Overview of available options.

* `--worker <WORKER>`: *(Required)* Hostname of the worker that the master can use to access the worker. If the master can resolve the worker's hostname, one may specify\
  `--worker $(hostname -f)`
* `--master <MASTER>`: *(Required)* Hostname or local IP address of the CryoSPARC master computer
* `--port <PORT>`: Port on which the master node is running (*default* `39000`)
* `--update`: Update an existing worker configuration instead of registering a new one
* `--sshstr <SSHSTR>`: SSH-login string for the master to use to send commands to the workers (*default* `$(whoami)@<WORKER>`)
* `--nogpu`: Registers a worker without any GPUs installed or with GPU access disabled
* `--gpus <GPUS>`: Comma-separated list of GPU slots indexes that this worker has access to, starting with `0`. e.g., to only enable the last two GPUs on a 4-GPU machine, enter\
  `--gpus 2,3`
* `--nossd`: If specified, this worker does not have access to a Solid-State Drive (SSD) to use for caching particle data
* `--ssdpath <SSDPATH>`: Path to the location of the mounted SSD drive on the file system to use for caching. e.g., `--ssdpath /scratch/cryosparc_cache`
* `--ssdquota <SSDQUOTA>`: The maximum amount of space on the SSD that CryoSPARC uses on this worker, in megabytes (MB). Workers automatically remove older cache files to remain below this quota
* `--ssdreserve <SSDRESERVE>`: The amount of space to initially reserve on the SSD for this worker
* `--lane <LANE>`: Which of CryoSPARC's scheduler lanes to add this worker to. Use with\
  `--newlane` to create a new lane with the unique name `<LANE>`
* `--newlane`: Create a new scheduler lane to use for this worker instead of using an existing one
* `--rams <NUM_RAMS>`: an integer representing the number of 8GB system RAM slots to make available for allocation to CryoSPARC jobs. Allocation is based on job type-specific estimates and does not enforce a limit on the amount of RAM that a running job will eventually use
* `--cpus <NUM_CPUS>`: an integer representing the number of CPU cores to make available for allocation to CryoSPARC jobs
* `--monitor_port <MONITOR_PORT>`: Not used

### `cryosparcw gpulist`

Lists which GPUs the CryoSPARC worker processes have access to.

```
$ cryosparcw gpulist
  Detected 4 CUDA devices.

   id           pci-bus  name
   ---------------------------------------------------------------
       0      0000:02:00.0  Quadro GP100
       1      0000:84:00.0  Quadro GP100
       2      0000:83:00.0  GeForce GTX 1080 Ti
       3      0000:03:00.0  GeForce GTX TITAN X
   ---------------------------------------------------------------
```

Use this to verify that the worker is installed correctly

### `cryosparcw ipython`

Starts an [**ipython**](https://ipython.readthedocs.io/en/stable/) shell in CryoSPARC's worker environment.

### `cryosparcw patch`

Install a patch previously downloaded on the master node with `cryosparcm patch --download` that is copied to the worker installation directory.

[See `cryosparcm patch` documentation.](/setup-configuration-and-management/management-and-monitoring-4.7/cryosparcm-4.7#cryosparcm-patch)

### `cryosparcw update`

Used for manual worker updates. See [**Software Updates**](/setup-configuration-and-management/software-updates#manual-cluster-updates) for details

### `cryosparcw newcuda <path>` *(CryoSPARC v4.3 and older)*

{% hint style="info" %}
CryoSPARC versions v4.4.0 and newer include a bundled CUDA Toolkit and no longer have this command.
{% endhint %}

Specifies a new path to the CUDA installation to use. Example usage:

```bash
cryosparcw newcuda /usr/local/cuda-11.0
```

###

{% content-ref url="/pages/-M7DHIJwDbQvhcKm7avb" %}
[Software Updates and Patches](/setup-configuration-and-management/software-updates)
{% endcontent-ref %}


# Management and Monitoring (v5.0+)

Instructions for accessing and working in the CryoSPARC command line.

## Environment variables

Specify additional environment variables in the configuration files to augment CryoSPARC's low-level behaviour.

{% content-ref url="/pages/FE8HagwiwWP6ayaVxPTy" %}
[Environment Variables (v5.0+)](/setup-configuration-and-management/management-and-monitoring-v5.0/environment-variables-v5.0)
{% endcontent-ref %}

## cryosparcm, cryosparcm cli and cryosparcw references

Workstations or master nodes with a `cryosparc_master` installation have access to `cryosparcm`, CryoSPARC's built-in [command-line](https://en.wikipedia.org/wiki/Command-line_interface) utility for all administrative, management and advanced usage tasks.

{% content-ref url="/pages/U4loQEjnD54VcG9JSUTi" %}
[cryosparcm reference (v5.0+)](/setup-configuration-and-management/management-and-monitoring-v5.0/cryosparcm-reference-v5.0)
{% endcontent-ref %}

The `cryosparcm cli` command provides an extensive API for programmatically controlling cryoSPARC from the command-line.

{% content-ref url="/pages/wIIf9eVztQ9carW68Oe7" %}
[cryosparcm cli reference (v5.0+)](/setup-configuration-and-management/management-and-monitoring-v5.0/cryosparcm-cli-reference-v5.0)
{% endcontent-ref %}

Workstations or worker nodes with a `cryosparc_worker` installation have access to `cryosparcw`, a utility similar to `cryosparcm` for managing worker installations.

{% content-ref url="/pages/VFHjc1Ode6NjypOczENX" %}
[cryosparcw reference (v5.0+)](/setup-configuration-and-management/management-and-monitoring-v5.0/cryosparcw-reference-v5.0)
{% endcontent-ref %}


# Environment Variables (v5.0+)

(Advanced) Specify additional environment variables in the configuration files to augment CryoSPARC's low-level behaviour.

## Environment variables

To set or change one of these environment settings, add a new line to one the `config.sh` files with the following format (substitute `VARIABLE` and `VALUE` as indicated in the next sections):

```bash
export VARIABLE="VALUE"
```

Or set the value based on a different environment variable provided by the system:

```bash
export VARIABLE="${OTHER_VARIABLE}"
```

Which `config.sh` file you use depends on which variable you need to change. The variables available for each file are described below.

### cryosparc\_master/config.sh

**Note:** Restart CryoSPARC with `cryosparcm restart` after changing this file.

<table><thead><tr><th width="299.546875">Variable</th><th width="300.28515625">Description</th><th width="198.7265625">Default Value</th></tr></thead><tbody><tr><td><code>CRYOSPARC_API_PROCS</code></td><td>Number of API service processes (1–8).</td><td><code>3</code></td></tr><tr><td><code>CRYOSPARC_AUTO_EXPORT_INSTANCE_CONFIG_DIR</code></td><td>The location which instance configurations will be automatically exported to.</td><td><code>$ROOT_DIR/run</code></td></tr><tr><td><code>CRYOSPARC_AUTO_EXPORT_INSTANCE_CONFIG_ENABLE</code></td><td>Enables automatic instance configuration exports.</td><td><code>true</code></td></tr><tr><td><code>CRYOSPARC_AUTO_EXPORT_INSTANCE_CONFIG_INTERVAL_HOURS</code></td><td>Interval for the automatic export of instance configurations.</td><td><code>1</code></td></tr><tr><td><code>CRYOSPARC_CLI_SKIP_ACCESS_CHECK</code></td><td>Set to <code>1</code> or <code>true</code> to disable filesystem permission checks for <code>cryosparcm</code> command-line arguments; may be used with network file systems where listed UNIX access permissions do not match the true available permissions.</td><td><code>false</code></td></tr><tr><td><code>CRYOSPARC_CLUSTER_JOB_MONITOR_INTERVAL</code></td><td>The amount of time to wait (in seconds) in between status updates for cluster jobs.</td><td><code>10</code></td></tr><tr><td><code>CRYOSPARC_CLUSTER_JOB_MONITOR_MAX_RETRIES</code></td><td>The maximum amount of retries for cluster job status updates.</td><td><code>1000000</code></td></tr><tr><td><code>CRYOSPARC_DB_CONNECTION_TIMEOUT_MS</code></td><td>The amount of time to wait (in milliseconds) for the database to respond when starting it.</td><td><code>20000</code></td></tr><tr><td><code>CRYOSPARC_DB_ENABLE_AUTH</code></td><td>Enables authentication for MongoDB database operations</td><td><code>true</code></td></tr><tr><td><code>CRYOSPARC_DB_MIN_SPACE_GB</code></td><td>Minimum free disk space required for MongoDB data directory.</td><td><code>5</code></td></tr><tr><td><code>CRYOSPARC_DISABLE_EXTERNAL_REQUESTS</code></td><td>Disables application requests to external HTTPS resources used for information modules in the homepage including EMPIAR, EMDB, Discuss, CryoSPARC tutorials, and the CryoSPARC changelog.</td><td><code>false</code></td></tr><tr><td><code>CRYOSPARC_FORCE_HOSTNAME</code></td><td>In master/worker or cluster modes, <code>cryosparcm</code> commands always run on the machine where <code>cryosparc_master</code> was initially installed. Set this to <code>true</code> to allow running on any machine</td><td><code>false</code></td></tr><tr><td><code>CRYOSPARC_FORCE_USER</code></td><td><code>cryosparcm</code> commands must be run by the same UNIX user account that owns the <code>cryosparc_master</code> installation. Set this to <code>true</code> to allow running <code>cryosparcm</code> from any user account</td><td><code>false</code></td></tr><tr><td><code>CRYOSPARC_HEARTBEAT_SECONDS</code></td><td>CryoSPARC jobs running on worker nodes regularly report their status to the master command server to indicate that they are still running and active. If a job fails to report for more than this number of seconds (e.g., due to stalling, a slow network or a silent error), CryoSPARC marks the job as failed. Increase for very busy/low-resource worker nodes or slow/unreliable connections between the master and worker nodes. Increasing may reduce heartbeat-related job failures.</td><td><code>180</code></td></tr><tr><td><code>CRYOSPARC_IGNORE_HIDDEN_FILES</code></td><td>When checking that job directory is empty before running a job, set to <code>true</code> to ignore any filenames that being with <code>.</code>. Use this if your file system or OS automatically creates hidden files</td><td><code>false</code></td></tr><tr><td><code>CRYOSPARC_IGNORE_PORT_CONFLICTS</code></td><td>Ignore port conflicts when binding service ports.</td><td><code>false</code></td></tr><tr><td><code>CRYOSPARC_INSECURE</code></td><td>Disables HTTPS secure connections and enables HTTP insecure connections, for connections to CryoSPARC license servers.</td><td><code>false</code></td></tr><tr><td><code>CRYOSPARC_IO_URING</code></td><td>Set to <code>false</code> to force-disable io_uring for fast disk read operations during jobs. May be required for some older operating systems or systems with an incorrect io_uring implementation.</td><td><code>true</code></td></tr><tr><td><code>CRYOSPARC_LICENSE_SERVER_ADDR</code></td><td>Override CryoSPARC license server address. Use for systems that require access through a proxy.</td><td><code>https://get.cryosparc.com</code></td></tr><tr><td><code>CRYOSPARC_LIVE_RETRY_ATTEMPTS</code></td><td>Number of attempts to retry a failed worker job in CryoSPARC Live. If a worker job fails more times than this value, the job will not be automatically restarted.</td><td><code>3</code></td></tr><tr><td><code>CRYOSPARC_LIVE_RETRY_TIMEOUT_SECONDS</code></td><td>Length of time to wait before restarting a failed worker job in CryoSPARC Live</td><td><code>30</code></td></tr><tr><td><code>CRYOSPARC_MONGO_CACHE_GB</code></td><td>How much RAM to allocate for MongoDB database query cache, in GB</td><td><code>4</code></td></tr><tr><td><code>CRYOSPARC_MOTION_CORRECTION_CPUS_PER_GPU</code></td><td>How many CPU cores to allocate for each GPU used in Patch, Full-frame and Local Motion Correction jobs</td><td><code>6</code></td></tr><tr><td><code>CRYOSPARC_MOTION_CORRECTION_RAM_MB_PER_GPU</code></td><td>How much RAM (in MB) to allocate for each GPU used in Patch, Full-frame and Local Motion Correction jobs</td><td><code>15000</code></td></tr><tr><td><code>CRYOSPARC_PROJECT_DIR_PREFIX</code></td><td>The prefix to add to a project directory name on the filesystem when it is created.</td><td><code>CS-</code></td></tr><tr><td><code>CRYOSPARC_RETAIN_QUEUE_POSITION_SECONDS</code></td><td>When clearing and re-queuing a job, this value is the window of time for which the job will retain its position in the CryoSPARC instance job queue. i.e. The job will still count as being scheduled from when it was first queued for <code>CRYOSPARC_RETAIN_QUEUE_POSITION_SECONDS</code> long since it was cleared. After the window passes without re-queueing, the job will lose its position in the queue, and the next time it is queued it will schedule as normal.</td><td><code>0</code></td></tr><tr><td><code>CRYOSPARC_SLACK_WEBHOOK_URL</code></td><td>If set, CryoSPARC makes an HTTP request to this URL each time a job's status changes with some job metadata encoded in JSON</td><td>-</td></tr><tr><td><code>CRYOSPARC_SSD_CACHE_LIFETIME_DAYS</code></td><td>Whenever a job requires SSD cache, it automatically checks for and removes files that haven't been accessed in more than the number of days specified by this variable. <strong>Note:</strong> Files may remain on the SSD for longer than this amount since they only get cleaned up when a job runs. Files may remain on the SSD for shorter than this amount if they are not in use by an active job and additional space is needed for caching of other particles.</td><td><code>30</code></td></tr><tr><td><code>REQUESTS_CA_BUNDLE</code></td><td>May be required for connection to the CryoSPARC license verification server on systems with outdated Certificate-Authority files or that filter HTTPS requests through a proxy. Specify a path to a file or folder that contains the certificates.</td><td>-</td></tr></tbody></table>

**Note:** This file includes the following environment variables that are specified at installation time:

* `CRYOSPARC_LICENSE_ID`
* `CRYOSPARC_MASTER_HOSTNAME`
* `CRYOSPARC_DB_PATH`
* `CRYOSPARC_BASE_PORT`

### cryosparc\_worker/config.sh

No restart is required after changing this file.

<table><thead><tr><th width="300.3671875">Variable</th><th width="300.34765625">Description</th><th width="186.1015625">Default Value</th></tr></thead><tbody><tr><td><code>CRYOSPARC_CACHE_COPY_BUFFER_SIZE_MB</code></td><td>Size of the buffer used when copying files to the CryoSPARC cache location.</td><td><code>8</code></td></tr><tr><td><code>CRYOSPARC_CACHE_LOCK_STRATEGY</code></td><td>Distributed locking strategy to use when multiple running jobs access the SSD cache simultaneously. Set to <code>file</code> to use a file-system POSIX lock. Set to <code>master</code> to use the CryoSPARC master as a broker. Use the <code>master</code> strategy for caches on distributed file systems such as GPFS and BeeGFS where POSIX locks are disabled or unavailable.</td><td><code>master</code></td></tr><tr><td><code>CRYOSPARC_CACHE_MAX_ATTEMPTS</code></td><td>Number of attempts to retry using the CryoSPARC cache. The job will fail if the number of attempts exceeds this value.</td><td><code>1</code></td></tr><tr><td><code>CRYOSPARC_CACHE_NUM_THREADS</code></td><td>Number of threads to use during caching when copying particle <code>.mrc</code> files from the project directory to the cache directory. Set to <code>1</code> to disable threading and copy files sequentially. See <a data-mention href="/pages/-MNeFvdldyAikdOWQ4K3#leveraging-multiple-threads-to-copy-particles">/pages/-MNeFvdldyAikdOWQ4K3#leveraging-multiple-threads-to-copy-particles</a> for details.</td><td><code>2</code></td></tr><tr><td><code>CRYOSPARC_CLI_SKIP_ACCESS_CHECK</code></td><td>Set to <code>1</code> or <code>true</code> to disable filesystem permission checks for <code>cryosparcm</code> command-line arguments; may be used with network file systems where listed UNIX access permissions do not match the true available permissions.</td><td><code>false</code></td></tr><tr><td><code>CRYOSPARC_IO_FD_LIMIT</code></td><td>Limit of how many open file descriptors CryoSPARC can create when running jobs.</td><td>-</td></tr><tr><td><code>CRYOSPARC_IO_URING</code></td><td>Set to <code>false</code> to force-disable io_uring for fast disk read operations during jobs. May be required for some older operating systems or systems with an incorrect io_uring implementation.</td><td><code>true</code></td></tr><tr><td><code>CRYOSPARC_NO_PAGELOCK</code></td><td>By default, CryoSPARC uses the CUDA driver's <code>pagelocked_empty</code> function to allocate GPU memory. This causes CUDA-driver errors on some systems. Set this variable to <code>true</code> to use <code>numpy.empty</code></td><td><code>false</code></td></tr><tr><td><code>CRYOSPARC_SSD_PATH</code></td><td>Set this variable to override the SSD cache path provided when you installed the worker. Useful if the SSD cache path is generated by your cluster as an environment variable when the job is scheduled. <strong>Important</strong>: The worker must be connected with a stub SSD path for this to take effect, e.g., <code>"cache_path": "/tmp"</code> in <code>cluster_config.json</code></td><td>-</td></tr></tbody></table>

**Note:** This file includes the following environment variables that are specified at installation time:

* `CRYOSPARC_LICENSE_ID`


# cryosparcm reference (v5.0+)

How to use the cryosparcm utility for starting and stopping the CryoSPARC instance, checking status or logs, managing users and using CryoSPARC's command-line interface.

## Access the CryoSPARC command line utility, `cryosparcm`

The CryoSPARC master node hosts the web server and manages job resource allocation.

Workstations or master nodes with a `cryosparc_master` installation have access to `cryosparcm`, CryoSPARC's built-in [command-line](https://en.wikipedia.org/wiki/Command-line_interface) utility for all administrative, management and advanced usage tasks.

To use it, log into the machine onto which [CryoSPARC was installed](https://guide.cryosparc.com/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc). Open a Terminal running a shell (such as `bash`) and enter any of the commands described below.

{% hint style="info" %}
If CryoSPARC was installed without adding CryoSPARC’s `bin` path to the shell’s path configuration, navigate to the `cryosparc_master` installation directory and run `./bin/cryosparcm` instead of `cryosparcm`.
{% endhint %}

For help with a specific command, run:

```bash
cryosparcm COMMAND --help
```

For example, for help starting CryoSPARC, run:

```bash
cryosparcm start --help
```

### `cryosparcm`

**Usage**:

```bash
$ cryosparcm [OPTIONS] COMMAND [ARGS]...
```

**Options**:

* `--install-completion`: Install completion for the current shell.
* `--show-completion`: Show completion for the current shell, to copy it or customize the installation.
* `--help`: Show this message and exit.

All available `cryosparcm` commands are listed and documented in the sections below.

## Instance Status and Management

Always run instance management commands in this section from the UNIX user account that owns the CryoSPARC installation, and always on the same machine on the network that `cryosparc_master` was installed on. If these conditions are not met, you may see the following message:

```bash
$ cryosparcm status
────────────────────────────────────────────────────────────────────────────────
CryoSPARC System master node installed at
/home/cryosparcuser/cryosparc_master
Current CryoSPARC version: develop
────────────────────────────────────────────────────────────────────────────────

✕ UnauthorizedException: This command must run on the CryoSPARC master host, but
there is a mismatch between the $CRYOSPARC_MASTER_HOSTNAME definition
(example.edu) and the configured hostname of this host (gpu.example.xyz).

If, and only if, the command ran on the CryoSPARC master host, but the host is
configured with a different hostname, consider the appropriate intervention for
your circumstances:

 1 Ensure $CRYOSPARC_MASTER_HOSTNAME and the output of command hostname -f
   match. CRYOSPARC_MASTER_HOSTNAME may be defined inside
   cryosparc_master/config.sh. Restart CryoSPARC after this change.
 2 Or: re-run this command with the environment variable
   CRYOSPARC_FORCE_HOSTNAME="true". This setting bypasses an important safety
   check and may disrupt CryoSPARC function if used inappropriately.
```

You can temporarily force `cryosparcm` to ignore the current hostname or user by specifying the `CRYOSPARC_FORCE_HOSTNAME` or `CRYOSPARC_FORCE_USER` variables just before calling the command:

```bash
$ CRYOSPARC_FORCE_HOSTNAME=true cryosparcm status
```

If you see the above error message, but the hostname it reports is incorrect (i.e., the hostname specified in the error message is actually the same host, just a different identifier), you can set `CRYOSPARC_MASTER_HOSTNAME` in `cryosparc_master/config.sh` to the correct hostname. You can also set `CRYOSPARC_FORCE_HOSTNAME` or `CRYOSPARC_FORCE_USER` in this file to permanently suppress this message.

### `cryosparcm status`

Show CryoSPARC system status, including the status of all CryoSPARC processes (`database`, `app`, `api`, etc.) and show configuration environment variables.

**Usage**:

```
$ cryosparcm status [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

### `cryosparcm version`

Show CryoSPARC version.

**Usage**:

```bash
$ cryosparcm version [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

### `cryosparcm start`

Start CryoSPARC or one of its services.

All processes start in the background, including all services and the web interface; processes will continue running after the terminal is closed. To stop, use `cryosparcm stop`. Provide an optional service name to only start that specific service.

**Usage**:

```
$ cryosparcm start [OPTIONS] [SERVICE]:[app|database|cache|api|scheduler|command_vis|app_api]
```

**Arguments**:

* `[SERVICE]:[app|database|cache|api|scheduler|command_vis|app_api]`

**Options**:

* `--systemd / --no-systemd`: \[default: no-systemd]
* `--startup / --no-startup`: \[default: startup]
* `--app / --no-app`: \[default: app]
* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

### `cryosparcm stop`

Stop CryoSPARC or one of its services.

Provide an optional service name to only start that specific service.

**Usage**:

```
$ cryosparcm stop [OPTIONS] [SERVICE]:[app|database|cache|api|scheduler|command_vis|app_api]
```

**Arguments**:

* `[SERVICE]:[app|database|cache|api|scheduler|command_vis|app_api]`

**Options**:

* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

### `cryosparcm restart`

Stop and start CryoSPARC or one of its services.

**Usage**:

```
$ cryosparcm restart [OPTIONS] [SERVICE]:[app|database|cache|api|scheduler|command_vis|app_api]
```

**Arguments**:

* `[SERVICE]:[app|database|cache|api|scheduler|command_vis|app_api]`

**Options**:

* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

### `cryosparcm maintenancemode`

Enable, disable or check maintenance mode. While enabled, prevents queued jobs from running while allowing running jobs to finish. Improves user experience user experience while CryoSPARC is undergoing maintenance, for example during restart, patch, or update.

See [Guide: Maintenance Mode and Configurable User Facing Messages](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/guide-maintenance-mode-and-configurable-user-facing-messages) for full details.

**Usage**:

```
$ cryosparcm maintenancemode [OPTIONS] COMMAND:{status|on|off}
```

**Arguments**:

* `COMMAND:{status|on|off}`: \[required]

**Options**:

* `--help`: Show this message and exit.

### `cryosparcm resources`

Print a formatted table of available scheduler targets and their properties.

**Usage**:

```
$ cryosparcm resources [OPTIONS] [LANE_NAME]
```

**Arguments**:

* `[LANE_NAME]`: Only show target information for a specific lane

**Options**:

* `--help`: Show this message and exit.

### `cryosparcm changeport`

Change instance base port.

**Usage**:

```
$ cryosparcm changeport [OPTIONS] PORT
```

**Arguments**:

* `PORT`: \[required]

**Options**:

* `-y, --yes`: Confirm without prompting
* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

### `cryosparcm asset-stats`

Show asset storage statistics.

**Usage**:

```
$ cryosparcm asset-stats [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

### `cryosparcm recover`

Restore instance configuration and recover projects from an exported instance configuration file. Should only run when the database has no projects. For full instructions, see [Instance Recovery](/setup-configuration-and-management/software-system-guides/guide-instance-recovery-v5.0).

**Usage**:

```
$ cryosparcm recover [OPTIONS]
```

**Options**:

* `-f, --file FILE`: Path to input file \[required]
* `--claim-project-ownership`: Take over projects locked to other instances
* `-y, --yes`: Confirm without prompting
* `--help`: Show this message and exit.

## Instance Setup

### `cryosparcm update`

Install the latest CryoSPARC update. See [**Software Updates**](https://guide.cryosparc.com/setup-configuration-and-management/software-updates) for full details.

**Usage**:

```
$ cryosparcm update [OPTIONS]
```

**Options**:

* `--version TEXT`: Version to update to \[default: latest]
* `--list`: List available versions
* `--check`: Check for update
* `--download`: Download only
* `--install`: Install previous download
* `--force`: Force install the latest or specified version
* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

### `cryosparcm patch`

Download and install the latest patch for this version.

**Usage**:

```
$ cryosparcm patch [OPTIONS] [PATCH_NAME]
```

**Arguments**:

* `[PATCH_NAME]`: Name of patch to download

**Options**:

* `--check`: Check to see if a patch is available
* `--download`: Download patches for manual installation
* `--install`: Manually install a downloaded patch file
* `-f, --force`: Force install or re-install latest patch
* `-y, --yes`: Confirm patch installation without prompt
* `--help`: Show this message and exit.

Frequently used commands:

* `cryosparcm patch`: Automatically install the latest patches on workstations or master node and connected workers. *Not recommended for clusters: Use the* `--download` *and* `--install` *flags instead.*
* `cryosparcm patch --force`: Reinstall the latest patches in case something went wrong with a previous attempt
* `cryosparcm patch --check`: Show information about the latest patches without installing
* `cryosparcm patch --download`: Download the latest patches without installing them. Follow the resulting instructions to install the master and worker patches
* `cryosparcm patch --install`: Run this command immediately after a `--download` to install the patch on the master node.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

### `cryosparcm worker`

Worker management commands.

{% hint style="warning" %}
Ensure CryoSPARC is running before running worker management commands.
{% endhint %}

**Usage**:

```
$ cryosparcm worker [OPTIONS] COMMAND [ARGS]...
```

**Options**:

* `--help`: Show this message and exit.

**Commands**:

* `update`: Install a cryosparc worker update on all connected workers.
* `patch`: Install a cryosparc worker patch on all connected workers.
* `connect`: Connect a worker node that jobs can be scheduled on.
* `disconnect`: Remove a worker node from the scheduler.

#### `cryosparcm worker update`

Install a cryosparc worker update on all workers or the given worker.

**Usage**:

**Arguments**:

* `[WORKER]`: Target name. Applies to all targets if not specified

**Options**:

* `--file FILE`: \[default: cryosparc\_worker.tar.gz]
* `--force`
* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

#### `cryosparcm worker patch`

Install a cryosparc worker patch on all workers or the given worker.

**Usage**:

```
$ cryosparcm worker patch [OPTIONS] [WORKER]

```

**Arguments**:

* `[WORKER]`: Target name. Applies to all targets if not specified

**Options**:

* `--file FILE`: \[default: cryosparc\_worker\_patch.tar.gz]
* `--force`
* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

#### `cryosparcm worker connect`

Connect a worker node that jobs can be scheduled on, or update an existing worker configuration. Similar to [cryosparcw reference (v5.0+)](/setup-configuration-and-management/management-and-monitoring-v5.0/cryosparcw-reference-v5.0#cryosparcw-connect).

**Usage**:

```
$ cryosparcm worker connect [OPTIONS]
```

**Options**:

* `--path TEXT`: Path to cryosparc\_worker folder \[required]
* `--worker TEXT`: Name of worker. Defaults to $(hostname) if not specified.
* `--lane TEXT`: Scheduler lane for worker (create if does not exist). \[default: default]
* `--sshstr TEXT`: SSH login string to access worker, required if the worker's hostname or UNIX user differs when connecting from master, or to specify additional SSH flags. Defaults to "$(whoami)@worker".
* `--cpus INTEGER RANGE`: Number of CPU cores to enable for jobs. Enable all cores if not specified. \[x>=1]
* `--rams INTEGER RANGE`: Number of 8GiB RAM slots to enable for jobs. Enable all RAM if not specified. \[x>=1]
* `--gpus TEXT`: Comma-separated list of GPU device IDs, e.g., '0,1,2'. Selects all GPUs if not specified. Cannot be specified with --no-gpu.
* `--gpu / --no-gpu`: Do not attempt to select any GPUs. Don't specify both --no-gpu and --gpus flag. \[default: gpu]
* `--ssdpath TEXT`: Local SSD scratch path. Strongly recommended.
* `--ssdquota INTEGER`: Maximum amount of SSD space to use for caching, in megabytes (MB).
* `--ssdreserve INTEGER`: Minimum amount free space to leave on the SSD, in megabytes (MB). \[default: 10000]
* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

#### `cryosparcm worker disconnect`

Remove a worker node from the scheduler.

**Usage**:

```
$ cryosparcm worker disconnect [OPTIONS]
```

**Options**:

* `--worker TEXT`: Name of worker. Defaults to $(hostname) if not specified. \[required]
* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

### `cryosparcm cluster`

Cluster management commands. See the [Download and Installation](https://guide.cryosparc.com/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc) page for full details.

{% hint style="warning" %}
Ensure CryoSPARC is running before running cluster management commands.
{% endhint %}

**Usage**:

```
$ cryosparcm cluster [OPTIONS] COMMAND [ARGS]...
```

**Options**:

* `--help`: Show this message and exit.

**Commands**:

* `connect`: Create or update a cluster with...
* `dump`: Write cluster configuration and script to...
* `validate`
* `remove`: Remove a cluster from the scheduler.
* `example`: Write example cluster configuration and...

#### `cryosparcm cluster example`

Write example cluster configuration (`cluster_info.json`) and script (`cluster_script.sh`) to a directory.

Examples are available for [Portable Batch System](https://en.wikipedia.org/wiki/Portable_Batch_System) (`cryosparcm cluster example pbs`) and [SLURM](https://slurm.schedmd.com/documentation.html) (`cryosparcm cluster example slurm`) schedulers. Other systems are similar; run one of the two `cluster example` commands and modify the output files accordingly.

**Usage**:

```
$ cryosparcm cluster example [OPTIONS] TYPE:{pbs|slurm}
```

**Arguments**:

* `TYPE:{pbs|slurm}`: Any cluster scheduler is supported but may require a custom submission script. \[required]

**Options**:

* `-o, --output-dir DIRECTORY`: Path to output directory \[default: .]
* `--help`: Show this message and exit.

#### `cryosparcm cluster connect`

Create or update a cluster with `cluster_info.json` and `cluster_script.sh`.

**Usage**:

```
$ cryosparcm cluster connect [OPTIONS]
```

**Options**:

* `--info FILE`: \[default: cluster\_info.json]
* `--script FILE`: \[default: cluster\_script.sh]
* `--help`: Show this message and exit.

#### `cryosparcm cluster dump`

Write cluster configuration and script to a directory.

**Usage**:

```
$ cryosparcm cluster dump [OPTIONS] NAME
```

**Arguments**:

* `NAME`: Cluster target name \[required]

**Options**:

* `o, --output-dir DIRECTORY`: Path to output directory \[default: .]
* `--help`: Show this message and exit.

#### `cryosparcm cluster remove`

Remove a cluster from the scheduler.

**Usage**:

```
$ cryosparcm cluster remove [OPTIONS] NAME
```

**Arguments**:

* `NAME`: Cluster target name \[required]

**Options**:

* `--help`: Show this message and exit.

### `cryosparcm deps`

Install Python and external dependencies. Specify `--force` to install even if they haven't changed.

**Usage**:

```
$ cryosparcm deps [OPTIONS]
```

**Options**:

* `--force`
* `--help`: Show this message and exit.

### `cryosparcm test`

Verifies the instance has been correctly installed by running several tests. Provides a report upon completion. For more information, see [Guide: Installation Testing with cryosparcm test](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/guide-installation-testing-with-cryosparcm-test).

{% hint style="warning" %}
Ensure CryoSPARC is running before running worker management commands.
{% endhint %}

**Usage**:

```
$ cryosparcm test [OPTIONS] COMMAND [ARGS]...
```

**Options**:

* `--help`: Show this message and exit.

**Commands**:

* `license`: Verify that your CryoSPARC license is valid
* `install`: Test all installation components
* `i`: Alias for install
* `workers`: Test worker installation
* `w`: Alias for workers

#### `cryosparcm test license`

Verify that your CryoSPARC license is valid and that CryoSPARC can verify job runs with the license server at [get.cryosparc.com](http://get.cryosparc.com/).

**Usage**:

```
$ cryosparcm test license [OPTIONS]
```

**Options**:

* `-l, --long`
* `--help`: Show this message and exit.

#### `cryosparcm test install`

Tests the core installation components of CryoSPARC (HTTP connections, licensing, workers, etc.) that are required to start running jobs. Provides information on the status of the CryoSPARC instance (e.g., which version is running, whether a patch is available, etc.).

**Usage**:

```
$ cryosparcm test install [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

#### `cryosparcm test i`

Alias for `cryosparcm test install`

#### `cryosparcm test workers`

Test workers by running validation jobs in the specified project.

**Usage**:

```
$ cryosparcm test workers [OPTIONS] PROJECT
```

**Arguments**:

* `PROJECT`: \[required]

**Options**:

* `--test [all|launch|ssd|gpu]`: Specify either the launch, ssd or gpu test \[required]
* `-t, --target TEXT`: Specify one or more targets to run tests on, tests all if not specified
* `--test-pytorch / --no-test-pytorch`: Test if worker(s) can launch PyTorch jobs on all enabled GPUs \[default: no-test-pytorch]
* `--help`: Show this message and exit.

#### `cryosparcm test w`

Alias for `cryosparcm test workers`

## Logs

### `cryosparcm log`

Show a service output log from the most recent entries. The log is live-updated while the command-line remains open and new data is added to the log. To stop live updates and return to the shell, press `ctrl C` on your keyboard followed by `q`.

**Usage**:

```
$ cryosparcm log [OPTIONS] SERVICE:{app|database|cache|api|scheduler|command_vis|app_api|supervisord}
```

**Arguments**:

* `SERVICE:{app|database|cache|api|scheduler|command_vis|app_api|supervisord}`: \[required]

**Options**:

* `--help`: Show this message and exit.

To save the full log, redirect the output to a file. Example:

```bash
cryosparcm log api > api.log
```

To show only the last x*x* lines of the log, pipe to `tail`. For example, to see the last 1000 lines of the log:

Copy

```
cryosparcm log api | tail -n 1000
```

### `cryosparcm filterlog`

Show a filtered service output log.

**Usage**:

```
$ cryosparcm filterlog [OPTIONS] SERVICE:{app|database|cache|api|scheduler|command_vis|app_api|supervisord}

```

**Arguments**:

* `SERVICE:{app|database|cache|api|scheduler|command_vis|app_api|supervisord}`: \[required]

**Options**:

* `-d, --days INTEGER RANGE`: Show logs within previous N days \[default: 30; x>=0]
* `-D, --date [%Y-%m-%d]`: Show logs on date
* `-m, --max-lines INTEGER RANGE`: Max lines to show \[x>=0]
* `-t, --tail`: Continuously follow this log
* `--help`: Show this message and exit.

Note that only `database`, `api`, `scheduler` and `command_vis` services support date and days filters.

### `cryosparcm snaplogs`

Create an archive with all current master service logs.

**Usage**:

```
$ cryosparcm snaplogs [OPTIONS]
```

**Options**:

* `-o, --output-dir DIRECTORY`: Path to output directory \[default: .]
* `--help`: Show this message and exit.

### `cryosparcm errorreport`

Generate a diagnostic information bundle. For more information, see [Guide: Download Error Reports](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/guide-download-error-reports).

**Usage**:

```
$ cryosparcm errorreport [OPTIONS]
```

**Options**:

* `-o, --output-dir DIRECTORY`: Path to output directory \[default: .]
* `-d, --days INTEGER RANGE`: Show logs within previous N days \[default: 30; x>=0]
* `-D, --date [%Y-%m-%d]`: Show logs on date
* `-m, --max-lines INTEGER RANGE`: Max lines to show \[x>=0]
* `--offline / --no-offline`: Skip database and worker data \[default: no-offline]
* `--skip-workers / --no-skip-workers`: Skip worker data \[default: no-skip-workers]
* `--help`: Show this message and exit.

### `cryosparcm get-workspace-report`

Download HTML workspace report from the CryoSPARC app.

**Usage**:

```
$ cryosparcm get-workspace-report [OPTIONS]
```

**Options**:

* `--project TEXT`: \[required]
* `--workspace TEXT`: \[required]
* `-o, --path PATH`: Path to output file or directory \[default: .]
* `--help`: Show this message and exit.

## User Management

Functions provided by these commands are also available from the web interface. For more details, see the [Admin Panel](https://guide.cryosparc.com/application-guide-v4.0+/admin-panel#user-management) guide.

### `cryosparcm user`

User management commands.

**Usage**:

```
$ cryosparcm user [OPTIONS] COMMAND [ARGS]...
```

**Options**:

* `--help`: Show this message and exit.

**Commands**:

* `list`: Show a table of available user accounts
* `exists`: Check if user exists
* `create`: Create a new user account
* `update`: Update user account information
* `resetpassword`: Reset a user account's password

#### `cryosparcm user list`

Show a table of available user accounts, including their names, email address and admin status.

**Usage**:

```
$ cryosparcm user list [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

#### `cryosparcm user exists`

Check if user exists. Command exits with status 0 if user exists, 1 otherwise. Shows an “exists” or “does not exist” message in either case.

**Usage**:

```
$ cryosparcm user exists [OPTIONS]
```

**Options**:

* `--email TEXT`: Email \[required]
* `--help`: Show this message and exit.

#### `cryosparcm user create`

Create a new user account. Call without arguments to create interactively.

**Usage**:

```
$ cryosparcm user create [OPTIONS]
```

**Options**:

* `--email TEXT`: Login email \[required]
* `--password TEXT`: Password \[required]
* `--username TEXT`: User name \[required]
* `--firstname TEXT`: First or given name \[required]
* `--lastname TEXT`: Last or surname \[required]
* `--role [user|admin]`: User role \[default: user]
* `--help`: Show this message and exit.

\<aside> 💡

If any required options are not specified, an input prompt will be provided.

\</aside>

#### `cryosparcm user update`

Update user account information and access, providing the email and password to verify. To change the password, use `cryosparcm users resetpassword`. Note that the user's email cannot be changed with this command.

**Usage**:

```
$ cryosparcm user update [OPTIONS]
```

**Options**:

* `--email TEXT`: Email \[required]
* `--password TEXT`: Password \[required]
* `--username TEXT`: New user name
* `--firstname TEXT`: New first or given name
* `--lastname TEXT`: New last or surname
* `--role [user|admin]`: New user role
* `--help`: Show this message and exit.

\<aside> 💡

If any required options are not specified, an input prompt will be provided.

\</aside>

Other than for the first user account created, new users do not have administrative privileges by default. After creating the first user account, other accounts can also be created [through the user interface](https://guide.cryosparc.com/application-guide-v4.0+/admin-panel#user-management) if preferred.

#### `cryosparcm user resetpassword`

Reset a user account's password.

**Usage**:

```
$ cryosparcm user resetpassword [OPTIONS]
```

**Options**:

* `--email TEXT`: Email \[required]
* `--password TEXT`: Password \[required]
* `--help`: Show this message and exit.

{% hint style="warning" %}
If any required options are not specified, an input prompt will be provided.
{% endhint %}

## Job Management

### `cryosparcm job`

Job management commands.

**Usage**:

```
$ cryosparcm job [OPTIONS] COMMAND [ARGS]...
```

**Options**:

* `--help`: Show this message and exit.

**Commands**:

* `status`: Show a summary of queued and active jobs
* `queue`: Queue a job
* `clear`: Clear a job
* `kill`: Kill a job
* `log`: Show job standard output and error log
* `events`: Show job event log

#### `cryosparcm job status`

Show a summary of queued and active jobs.

**Usage**:

```
$ cryosparcm job status [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

#### `cryosparcm job queue`

Queue a job. Specify either `--lane` or `--hostname`, except for interactive jobs. One or more `--gpu` options may be specified with `--hostname`.

**Usage**:

```
$ cryosparcm job queue [OPTIONS] PROJECT_UID JOB_UID
```

**Arguments**:

* `PROJECT_UID`: \[required]
* `JOB_UID`: \[required]

**Options**:

* `-l, --lane TEXT`: Scheduler lane to queue to
* `-h, --hostname TEXT`: Worker node to queue to
* `-g, --gpu INTEGER`: Specify one or more GPUs to queue to, `--hostname` must also be specified
* `--check-inputs-ready / --no-check-inputs-ready`: If disabled, job will run even if parent input jobs are incomplete \[default: check-inputs-ready]
* `--help`: Show this message and exit.

#### `cryosparcm job clear`

Clear a job.

**Usage**:

```
$ cryosparcm job clear [OPTIONS] PROJECT_UID JOB_UID
```

**Arguments**:

* `PROJECT_UID`: \[required]
* `JOB_UID`: \[required]

**Options**:

* `--help`: Show this message and exit.

#### `cryosparcm job kill`

Kill a job by its project UID and job UID.

**Usage**:

```
$ cryosparcm job kill [OPTIONS] PROJECT_UID JOB_UID
```

**Arguments**:

* `PROJECT_UID`: \[required]
* `JOB_UID`: \[required]

**Options**:

* `--help`: Show this message and exit.

#### `cryosparcm job log`

Show job standard output and error log.

**Usage**:

```
$ cryosparcm job log [OPTIONS] PROJECT_UID JOB_UID
```

**Arguments**:

* `PROJECT_UID`: \[required]
* `JOB_UID`: \[required]

**Options**:

* `--help`: Show this message and exit.

#### `cryosparcm job events`

Show job event log.

**Usage**:

```
$ cryosparcm job events [OPTIONS] PROJECT_UID JOB_UID
```

**Arguments**:

* `PROJECT_UID`: \[required]
* `JOB_UID`: \[required]

**Options**:

* `-c, --checkpoint INTEGER`: Show events from this checkpoint up until the next one. Shows all events if not provided. Specify -1 to show everything after the last checkpoint.
* `--help`: Show this message and exit.

## Database Management

Always run instance management commands in this section from the UNIX user account that owns the CryoSPARC installation, and always on the same machine on the network that `cryosparc_master` was installed on.

{% hint style="danger" %}
Note when using `cryosparcm database backup` and `cryosparcm database restore` commands:

Once CryoSPARC projects or jobs are created, deleted, or otherwise modified during or after the backup, a database restored from the resulting backup file will no longer be compatible with the modified project directories.

Use the `cryosparcm recover` command instead of the `backup`/`restore` commands to avoid this issue. See [Guide: Instance Recovery (v5.0+)](/setup-configuration-and-management/software-system-guides/guide-instance-recovery-v5.0) for details.
{% endhint %}

### `cryosparcm database`

Database management commands.

**Usage**:

```
$ cryosparcm database [OPTIONS] COMMAND [ARGS]...
```

**Options**:

* `--help`: Show this message and exit.

**Commands**:

* `check`: Check that the database is running correctly.
* `configure`: Prepare the database for running.
* `fixport`: Update expected database port.
* `backup`: Make a backup copy of the database.
* `restore`: Restore the database from a backup file.
* `compact`: Attempt to reduce database size.
* `export`: Export the contents of the database.
* `import`: Import a collection from an exported .json file
* `export-instance-config`: Export instance configuration and project information.
* `import-instance-config`: Import instance configuration from an exported file.

#### `cryosparcm database check`

Check that the database is running with the correct host configuration.

**Usage**:

```
$ cryosparcm database check [OPTIONS]
```

**Options**:

* `--quiet / --no-quiet`: \[default: no-quiet]
* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

#### `cryosparcm database configure`

Prepare the database for running. Automatically runs during `start`, manual invocation not typically required.

**Usage**:

```
$ cryosparcm database configure [OPTIONS]
```

**Options**:

* `-v, --verbose`
* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

#### `cryosparcm database fixport`

Update expected database port so that it matches the configured port following a change to `CRYOSPARC_BASE_PORT` in [config.sh](http://config.sh).

**Usage**:

```
$ cryosparcm database fixport [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

#### `cryosparcm database backup`

Make a backup copy of the database with `mongodump`.

By default, and saves the backup as an `.archive` file to the current working directory, with the current date and time in in the filename, e.g., `cryosparc_backup_2021_06_14_11h27.archive`

{% hint style="warning" %}
Do not allow the backup to fill up the filesystem on which the database is stored. If needed, specify a custom alternative path where the backup will be written.
{% endhint %}

**Usage**:

```
$ cryosparcm database backup [OPTIONS]
```

**Options**:

* `-o, --output PATH`: Path to output file or directory \[default: .]
* `-c, --collection TEXT`
* `--help`: Show this message and exit.

{% hint style="danger" %}
Once CryoSPARC projects or jobs are created, deleted, or otherwise modified during or after the backup, a database restored from the resulting backup file will no longer be compatible with the modified project directories.

Use the `cryosparcm recover` command instead of the `backup`/`restore` commands to avoid this issue. See [Guide: Instance Recovery (v5.0+)](/setup-configuration-and-management/software-system-guides/guide-instance-recovery-v5.0) for details.
{% endhint %}

{% hint style="warning" %}
CryoSPARC can be running when `cryosparcm database backup` is run, but the backup will impact the performance of your running database ([source](https://www.mongodb.com/docs/v3.6/tutorial/backup-and-restore-tools/#back-up-and-restore-with-mongodb-tools)).
{% endhint %}

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

#### `cryosparcm database restore`

Restore the database from a backup file.

{% hint style="danger" %}
A database backup becomes outdated and incompatible with project directories as soon as CryoSPARC projects or jobs are created, deleted or modified following a database backup. Do not restore an outdated database backup. Restoration of an outdated database backup and subsequent use with CryoSPARC is likely to corrupt CryoSPARC projects.

Use the `cryosparcm recover` command instead of the `backup`/`restore` commands to avoid this issue. See [Guide: Instance Recovery (v5.0+)](/setup-configuration-and-management/software-system-guides/guide-instance-recovery-v5.0) for details.
{% endhint %}

{% hint style="warning" %}
CryoSPARC must be [installed](https://guide.cryosparc.com/setup-configuration-and-management/how-to-download-install-and-configure) and but not started before running this command.
{% endhint %}

**Usage**:

```
$ cryosparcm database restore [OPTIONS]
```

**Options**:

* `-f, --file FILE`: Path to input file \[required]
* `-c, --collection TEXT`
* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

#### `cryosparcm database compact`

Attempt to reduce database size.

**Usage**:

```
$ cryosparcm database compact [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

#### `cryosparcm database export`

Export the contents of a database collection to a .json file.

**Usage**:

```
$ cryosparcm database export [OPTIONS] COLLECTION_NAME
```

**Arguments**:

* `COLLECTION_NAME`: \[required]

**Options**:

* `-o, --output-dir DIRECTORY`: Path to output directory \[default: .]
* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

#### `cryosparcm database import`

Import a collection from an exported .json file. Overwrites existing data in that collection.

**Usage**:

```
$ cryosparcm database import [OPTIONS] COLLECTION_NAME
```

**Arguments**:

* `COLLECTION_NAME`: \[required]

**Options**:

* `-f, --file FILE`: Path to input file \[required]
* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

#### `cryosparcm database export-instance-config`

Export instance configuration and project information to a .tar file.

**Usage**:

```
$ cryosparcm database export-instance-config [OPTIONS]
```

**Options**:

* `-o, --output-dir DIRECTORY`: Path to output directory \[default: .]
* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

#### `cryosparcm database import-instance-config`

Imports instance configuration from an exported instance configuration .tar file.

**Usage**:

```
$ cryosparcm database import-instance-config [OPTIONS]
```

**Options**:

* `-f, --file FILE`: Path to input file \[required]
* `-y, --yes`: Confirm without prompting
* `--help`: Show this message and exit.

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

## Command Line Utilities

### `cryosparcm env`

Export environment variables used by CryoSPARC.

**Usage**:

```
$ cryosparcm env [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

Run this command with `eval` to define the variables output by the `env` command.

```bash
eval $(cryosparcm env)
```

### `cryosparcm call`

Run any command with the CryoSPARC environment. For example:

```bash
cryosparcm call python -c "import sys; print(sys.path)"
```

**Usage**:

```
$ cryosparcm call [COMMAND] [ARGS]...
```

Equivalent to running `eval $(cryosparcm env)` followed by the command.

### `cryosparcm python`

Run python command with the CryoSPARC environment. For example:

```bash
cryosparcm python -c "import sys; print(sys.path)"
```

**Usage**:

```
$ cryosparcm python [OPTIONS] [ARGS]...
```

### `cryosparcm ipython`

Run an interactive python shell with the CryoSPARC environment.

**Usage**:

```
$ cryosparcm ipython [OPTIONS]
```

### `cryosparcm cli`

Interact with CryoSPARC from the command line with Python expressions. See cryosparcm cli reference for full details.

**Usage**:

```
$ cryosparcm cli [OPTIONS] EXPRESSION
```

**Arguments**:

* `EXPRESSION`: \[required] Python expression

**Options**:

* `--help`: Show this message and exit.

### `cryosparcm icli`

Interact with CryoSPARC from an ipython shell. Can use `api` object in Python commands. See cryosparcm cli reference for full details.

**Usage**:

```
$ cryosparcm icli [OPTIONS] [ARGS]...
```

### `cryosparcm downloadtest`

Download a test dataset. For use with the [T20S Introductory Tutorial](https://guide.cryosparc.com/guides-for-v3/cryo-em-data-processing-in-cryosparc-introductory-tutorial#t-20-s-tutorial) or Installation Testing or Performance benchmarking.

**Usage**:

```
$ cryosparcm downloadtest [OPTIONS]
```

**Options**:

* `-o, --output-dir DIRECTORY`: Path to output directory \[default: .]
* `--dataset TEXT`: Which dataset to download. One of '10025', '10305', or 'PERFORMANCE\_BENCHMARK\_DATA' \[default: 10025]
* `--help`: Show this message and exit.

### `cryosparcm mongo`

Start a mongo [shell](https://docs.mongodb.com/manual/mongo/) for CryoSPARC’s local MongoDB database service.

**Usage**:

```
$ cryosparcm mongo [OPTIONS] [ARGS]...
```

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

### `cryosparcm redis`

Start a [redis-cli](https://redis.io/docs/latest/develop/tools/cli/) prompt for command line access to the Redis cache service.

**Usage**:

```
$ cryosparcm redis [OPTIONS] [ARGS]...
```

{% hint style="warning" %}
**Always run this command from the same host and UNIX user account originally used to install CryoSPARC.**
{% endhint %}

## Deprecated Commands

These commands were present in CryoSPARC v4 and remain in CryoSPARC v5 but should no longer be used; instead, the corresponding commands above should be used.

#### `cryosparcm checkdb` (Deprecated)

Use [`cryosparcm database check`](#cryosparcm-database-check) instead.

**Usage**:

```
$ cryosparcm checkdb [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

#### `cryosparcm configuredb` (Deprecated)

Use [`cryosparcm database configure`](#cryosparcm-database-configure) instead.

**Usage**:

```
$ cryosparcm configuredb [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

#### `cryosparcm fixdbport` (Deprecated)

Use [`cryosparcm database fixport`](#cryosparcm-database-fixport) instead.

**Usage**:

```
$ cryosparcm fixdbport [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

#### `cryosparcm backup` (Deprecated)

Use [`cryosparcm database backup`](#cryosparcm-database-backup) instead.

{% hint style="danger" %}
Once CryoSPARC projects or jobs are created, deleted, or otherwise modified during or after the backup, a database restored from the resulting backup file will no longer be compatible with the modified project directories.

Use the `cryosparcm recover` command instead of the `backup`/`restore` commands to avoid this issue. See [Guide: Instance Recovery (v5.0+)](/setup-configuration-and-management/software-system-guides/guide-instance-recovery-v5.0) for details.
{% endhint %}

**Usage**:

```
$ cryosparcm backup [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

#### `cryosparcm restore` (Deprecated)

Use [`cryosparcm database restore`](#cryosparcm-database-restore) instead.

{% hint style="danger" %}
Once CryoSPARC projects or jobs are created, deleted, or otherwise modified during or after the backup, a database restored from the resulting backup file will no longer be compatible with the modified project directories.

Use the `cryosparcm recover` command instead of the `backup`/`restore` commands to avoid this issue. See [Guide: Instance Recovery (v5.0+)](/setup-configuration-and-management/software-system-guides/guide-instance-recovery-v5.0) for details.
{% endhint %}

**Usage**:

```
$ cryosparcm restore [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

#### `cryosparcm compact` (Deprecated)

Use [`cryosparcm database compact`](#cryosparcm-database-compact) instead.

**Usage**:

```
$ cryosparcm compact [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

#### `cryosparcm listusers` (Deprecated)

Use [`cryosparcm user list`](#cryosparcm-user-list) instead.

**Usage**:

```
$ cryosparcm listusers [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

#### `cryosparcm createuser` (Deprecated)

Use [`cryosparcm user create`](#cryosparcm-user-create) instead.

**Usage**:

```
$ cryosparcm createuser [OPTIONS]
```

**Options**:

* `--email TEXT`: \[required]
* `--password TEXT`: \[required]
* `--username TEXT`: \[required]
* `--firstname TEXT`: \[required]
* `--lastname TEXT`: \[required]
* `--role [user|admin]`: \[default: user]
* `--help`: Show this message and exit.

#### `cryosparcm updateuser` (Deprecated)

Use [`cryosparcm user update`](#cryosparcm-user-update) instead.

**Usage**:

```
$ cryosparcm updateuser [OPTIONS]
```

**Options**:

* `--email TEXT`: \[required]
* `--password TEXT`: \[required]
* `--username TEXT`
* `--firstname TEXT`
* `--lastname TEXT`
* `--admin [true|false]`
* `--help`: Show this message and exit.

#### `cryosparcm resetpassword` (Deprecated)

Use [`cryosparcm user resetpassword`](#cryosparcm-user-resetpassword) instead.

**Usage**:

```
$ cryosparcm resetpassword [OPTIONS]
```

**Options**:

* `--email TEXT`: \[required]
* `--password TEXT`: \[required]
* `--help`: Show this message and exit.

#### `cryosparcm jobstatus` (Deprecated)

Use [`cryosparcm job status`](#cryosparcm-job-status) instead.

**Usage**:

```
$ cryosparcm jobstatus [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

#### `cryosparcm joblog` (Deprecated)

Use [`cryosparcm job log`](#cryosparcm-job-log) instead.

**Usage**:

```
$ cryosparcm joblog [OPTIONS] PROJECT JOB
```

**Arguments**:

* `PROJECT`: \[required]
* `JOB`: \[required]

**Options**:

* `--help`: Show this message and exit.

#### `cryosparcm eventlog` (Deprecated)

Use [`cryosparcm job events`](#cryosparcm-job-events) instead.

**Usage**:

```
$ cryosparcm eventlog [OPTIONS] PROJECT JOB
```

**Arguments**:

* `PROJECT`: \[required]
* `JOB`: \[required]

**Options**:

* `--help`: Show this message and exit.

#### `cryosparcm licensestatus` (Deprecated)

Use [`cryosparcm test license`](#cryosparcm-test-license) instead.

**Usage**:

```
$ cryosparcm licensestatus [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

#### `cryosparcm cluster validate` (Deprecated)

Use [`cryosparcm test workers`](#cryosparcm-test-workers) instead.

**Usage**:

```
$ cryosparcm cluster validate [OPTIONS] NAME
```

**Arguments**:

* `NAME`: Cluster target name \[required]

**Options**:

* `--projects-dir TEXT`: Absolute path to projects directory \[required]
* `--help`: Show this message and exit.


# cryosparcm cli reference (v5.0+)

How to use CryoSPARC's low-level command line interface to perform actions that can be performed in the UI.

## Accessing the command line interface

CryoSPARC exposes nearly all actions that can be taken in the user interface through a command-line interface (CLI) for programatic operation. The CLI can be used to query information about jobs, projects, users, etc. from a CryoSPARC instance, or to take actions such as creating projects, queueing jobs, etc. Each CLI command is a [Python](https://www.python.org) expression, in the format `api.FUNCTION(...)` or `api.NAMESPACE.FUNCTION(...)` , where `...` is a list of function arguments.

The CLI can be called in three ways:

1. From `bash` or other shells, `cryosparcm cli "EXPRESSION"` can be used to easily execute a single command. For example:

   ```bash
   cryosparcm cli "api.jobs.enqueue('P3', 'J42', lane='default')"
   ```
2. Using CryoSPARC’s built-in interactive shell that is started by calling `cryosparc icli`. Once the interactive shell is started, Python code can be executed directly:

   ```python
   $ cryosparcm icli

   Connected to cryoem0.sbi:61002
   api, db, gfs ready to use

   In [1]: api.jobs.enqueue('P3', 'J42', lane='default')

   In [2]:
   ```
3. Using the standalone [`cryosparc-tools`](http://tools.cryosparc.com) python package in a Python script, for example:

   ```python
   from cryosparc.tools import CryoSPARC

   cs = CryoSPARC(...)
   cs.api.jobs.enqueue('P3', 'J42', lane='default')

   ```

## Available CLI commands

The CryoSPARC CLI exposes Python functions of the form `api.FUNCTION(...)` or `api.NAMESPACE.FUNCTION(...)` , where `...` is a list of function arguments.

All available `NAMESPACE` options (e.g., `projects`, `jobs`) and namespace-specific `FUNCTION` options (e.g., `find_one`, `create`) and their arguments are listed and documented in the `cryosparc-tools` documentation for the CryoSPARC API:

{% embed url="<https://tools.cryosparc.com/api/cryosparc.api.html>" %}

{% hint style="warning" %}
The format and arguments of `cryosparcm cli` commands may change from one CryoSPARC version to the next. The documentation for the API linked above will be updated with each CryoSPARC release.

In addition to the CLI described here, the [`cryosparc-tools` Python package](https://tools.cryosparc.com) provides an additional method for writing custom cryo-EM data processing scripts and workflow automation scripts. cryosparc-tools implements a subset of the commands that are available via the CLI.
{% endhint %}


# cryosparcw reference (v5.0+)

How to use the cryosparcw utility for managing CryoSPARC workers

## Worker management with `cryosparcw`

CryoSPARC worker nodes host CryoSPARC’s compute programs and libraries used to run jobs.

Workstations or worker nodes with a `cryosparc_worker` installation have access to `cryosparcw`, a built-in command-line utility to perform advanced usage tasks.

Run all commands in this section while logged into the workstation or worker nodes where the `cryosparc_worker` package is installed.

Navigate to the worker installation directory and run command starting with `bin/cryosparcw`. For example:

```bash
cd /path/to/cryosparc_worker
bin/cryosparcw info --gpu
```

You may add `cryosparcw` to the worker's `PATH` by adding a line like this to the `~/.bashrc` file, replacing `/path/to/cryosparc_worker` with the actual path to the worker installation directory:

```bash
export PATH=/path/to/cryosparc_worker/bin:$PATH
```

And restart your shell. Run the following command to verify that `PATH` was updated correctly:

```
which cryosparcw
```

And ensure the output is `/path/to/cryosparc_worker/bin/cryosparcw` .

For help with a specific command, run:

```bash
cryosparcw COMMAND --help
```

For example, for help showing worker info, run:

```bash
cryosparcw info --help
```

### `cryosparcw`

**Usage**:

```
  $ cryosparcw [OPTIONS] COMMAND [ARGS]...
```

**Options**:

* `--install-completion`: Install completion for the current shell.
* `--show-completion`: Show completion for the current shell, to copy it or customize the installation.
* `--help`: Show this message and exit.

All available `cryosparcw` commands are listed and documented below.

### `cryosparcw version`

Show CryoSPARC version.

**Usage**:

```
$ cryosparcw version [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

### `cryosparcw update`

Install a previously-downloaded CryoSPARC update at `cryosparc_worker/cryosparc_worker.tar.gz`

**Usage**:

```
$ cryosparcw update [OPTIONS]
```

**Options**:

* `--force`
* `--help`: Show this message and exit.

### `cryosparcw deps`

Install Python and external dependencies. Specify `--force` to install even if they haven't changed.

**Usage**:

```
$ cryosparcw deps [OPTIONS]
```

**Options**:

* `--force`
* `--help`: Show this message and exit.

### `cryosparcw patch`

Install a patch that was previously-downloaded on the master with `cryosparcm patch --download`.

**Usage**:

```
$ cryosparcw patch [OPTIONS]
```

**Options**:

* `-f, --file FILE`: Path to input file \[default: cryosparc\_worker/cryosparc\_worker\_patch.tar.gz]
* `--force`
* `--help`: Show this message and exit.

### `cryosparcw env`

Export environment variables used by CryoSPARC. Execute the output to activate, e.g., `eval $(cryosparcw env)`

**Usage**:

```
$ cryosparcw env [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

### `cryosparcw connect`

Connect a managed worker to CryoSPARC. Run this command on the worker that you wish to register with CryoSPARC, or whose existing registration you wish to update.

For full details, see [Connecting A Worker Node](https://guide.cryosparc.com/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc#connecting-a-worker-node).

**Usage**:

```
$ cryosparcw connect [OPTIONS]
```

**Options**:

* `--license TEXT`: CryoSPARC license ID read from installed configuration file (`config.sh`) if not specified. \[env var: CRYOSPARC\_LICENSE\_ID; required]
* `--master TEXT`: Hostname of machine running cryosparc\_master. Specify "localhost" for current host. \[env var: CRYOSPARC\_MASTER\_HOSTNAME; required]
* `--port INTEGER`: Base port for CryoSPARC services. Usually 61000 unless changed during installation. \[env var: CRYOSPARC\_BASE\_PORT; default: 61000]
* `--auth / --no-auth`: Enable master database authentication \[env var: CRYOSPARC\_DB\_ENABLE\_AUTH; default: auth]
* `--timeout INTEGER`: Database connection timeout, in milliseconds \[env var: CRYOSPARC\_DB\_CONNECTION\_TIMEOUT\_MS; default: 20000]
* `--worker TEXT`: Name of worker. Defaults to $(hostname) if not specified.
* `--lane TEXT`: Scheduler lane for worker (create if does not exist). \[default: default]
* `--sshstr TEXT`: SSH login string to access worker, required if the worker's hostname or UNIX user differs when connecting from master, or to specify additional SSH flags. Defaults to "$(whoami)@worker".
* `--cpus INTEGER RANGE`: Number of CPU cores to enable for jobs. Enable all cores if not specified. \[x>=1]
* `--rams INTEGER RANGE`: Number of 8GiB RAM slots to enable for jobs. Enable all RAM if not specified. \[x>=1]
* `--gpus TEXT`: Comma-separated list of GPU device IDs, e.g., '0,1,2'. Selects all GPUs if not specified. Cannot be specified with --no-gpu.
* `--gpu / --no-gpu`: Do not attempt to select any GPUs. Don't specify both --no-gpu and --gpus flag. \[default: gpu]
* `--ssdpath DIRECTORY`: Local SSD scratch path. Strongly recommended.
* `--ssdquota INTEGER`: Maximum amount of SSD space to use for caching, in megabytes (MB).
* `--ssdreserve INTEGER`: Minimum amount free space to leave on the SSD, in megabytes (MB). \[default: 10000]
* `--help`: Show this message and exit.

Example command to connect a worker to a new lane with the same name.

```bash
cryosparcw connect \
	--worker $(hostname -f) \
	--master csmaster.local \
	--port 61000 \
	--ssdpath /scratch/cryosparc_cache \
	--lane $(hostname -s)
```

See also [cryosparcm reference (v5.0+)](/setup-configuration-and-management/management-and-monitoring-v5.0/cryosparcm-reference-v5.0#cryosparcm-worker-connect).

### `cryosparcw disconnect`

Disconnect a worker node. You may also use [`cryosparcm worker disconnect`](/setup-configuration-and-management/management-and-monitoring-v5.0/cryosparcm-reference-v5.0#cryosparcm-worker-disconnect) if the worker installation directory is no longer accessible.

**Usage**:

```
$ cryosparcw disconnect [OPTIONS]
```

**Options**:

* `--license TEXT`: CryoSPARC license ID, read from installed configuration file ([config.sh](http://config.sh/)) if not specified. \[env var: CRYOSPARC\_LICENSE\_ID; required]
* `--master TEXT`: Hostname of machine running cryosparc\_master. Specify "localhost" for current host. \[env var: CRYOSPARC\_MASTER\_HOSTNAME; required]
* `--port INTEGER`: Base port for CryoSPARC services. Usually 61000 unless changed during installation. \[env var: CRYOSPARC\_BASE\_PORT; default: 61000]
* `--auth / --no-auth`: Enable master database authentication \[env var: CRYOSPARC\_DB\_ENABLE\_AUTH; default: auth]
* `--timeout INTEGER`: Database connection timeout, in milliseconds \[env var: CRYOSPARC\_DB\_CONNECTION\_TIMEOUT\_MS; default: 20000]
* `--worker TEXT`: Name of worker. Defaults to $(hostname) if not specified.
* `--help`: Show this message and exit.

### `cryosparcw info`

Show system information.

**Usage**:

```
$ cryosparcw info [OPTIONS]
```

**Options**:

* `--gpu / --no-gpu`: Include GPU info \[default: no-gpu]
* `--format [text|json]`: Output format \[default: text]
* `--help`: Show this message and exit.

### `cryosparcw gpulist`

Show information about available GPUs on this host.

**Usage**:

```
$ cryosparcw gpulist [OPTIONS]
```

**Options**:

* `--format [text|json]`: Output format \[default: text]
* `--help`: Show this message and exit.

Use to verify that the worker is installed correctly.

### `cryosparcw call`

Run any command with the CryoSPARC worker environment.

**Usage**:

```
$ cryosparcw call [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

### `cryosparcw python`

Run python command with the CryoSPARC worker environment.

**Usage**:

```
$ cryosparcw python [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.

### `cryosparcw ipython`

Run interactive python shell with the CryoSPARC worker environment.

**Usage**:

```
$ cryosparcw ipython [OPTIONS]
```

**Options**:

* `--help`: Show this message and exit.


# Software System Guides

CryoSPARC software management guides.

## Installing and Updating

{% content-ref url="/pages/fDNdMevkXwzDGwZkG4Bn" %}
[Guide: Updating to CryoSPARC v5](/setup-configuration-and-management/software-system-guides/guide-updating-to-cryosparc-v5)
{% endcontent-ref %}

{% content-ref url="/pages/GyWjZRmeDkKU1Zn2jWv4" %}
[Guide: Updating to CryoSPARC v4](/setup-configuration-and-management/software-system-guides/guide-updating-to-cryosparc-v4)
{% endcontent-ref %}

## Verify Installation, Performance and Troubleshooting

{% content-ref url="/pages/J11jkcODBUNnCrVFVYTr" %}
[Guide: Installation Testing with cryosparcm test](/setup-configuration-and-management/software-system-guides/guide-installation-testing-with-cryosparcm-test)
{% endcontent-ref %}

{% content-ref url="/pages/-MNhrOW1ytZ8-U3WkV0r" %}
[Guide: Verify CryoSPARC Installation with the Extensive Validation Job (v4.3+)](/setup-configuration-and-management/software-system-guides/tutorial-verify-cryosparc-installation-with-the-extensive-workflow-sysadmin-guide)
{% endcontent-ref %}

{% content-ref url="/pages/MGn7YPrLc7KM4Yyzi7ai" %}
[Guide: Verify CryoSPARC Installation with the Extensive Workflow (≤v4.2)](/setup-configuration-and-management/software-system-guides/tutorial-verify-cryosparc-installation-with-the-extensive-workflow-sysadmin-guide-1)
{% endcontent-ref %}

{% content-ref url="/pages/nov71O9rKSgfnnpFqRfs" %}
[Guide: Performance Benchmarking (v4.3+)](/setup-configuration-and-management/software-system-guides/guide-performance-benchmarking-v4.3)
{% endcontent-ref %}

{% content-ref url="/pages/TSuhnNBKoO8zPqocYGrn" %}
[Guide: Download Error Reports](/setup-configuration-and-management/software-system-guides/guide-download-error-reports)
{% endcontent-ref %}

## Instance Management

{% content-ref url="/pages/ExdUPe4xFh3gYepmHpCq" %}
[Guide: Maintenance Mode and Configurable User Facing Messages](/setup-configuration-and-management/software-system-guides/guide-maintenance-mode-and-configurable-user-facing-messages)
{% endcontent-ref %}

## User Management

{% content-ref url="/pages/-MNecrUQzW1CYEnJCepj" %}
[Guide: User Management](/setup-configuration-and-management/software-system-guides/tutorial-user-management)
{% endcontent-ref %}

{% content-ref url="/pages/-MckPoLFT9SUA459kEcN" %}
[Guide: Multi-user Unix Permissions and Data Access Control](/setup-configuration-and-management/software-system-guides/unix-permissions-and-data-access-control)
{% endcontent-ref %}

{% content-ref url="/pages/CJFjwLOpeqI6r91Mxkob" %}
[Guide: Lane Assignments and Restrictions](/setup-configuration-and-management/software-system-guides/guide-lane-assignments-and-restrictions)
{% endcontent-ref %}

## Job Queuing

{% content-ref url="/pages/-MNf1QrDa0ZRSzwp75Vz" %}
[Guide: Priority Job Queuing](/setup-configuration-and-management/software-system-guides/tutorial-priority-job-queuing)
{% endcontent-ref %}

{% content-ref url="/pages/5zikU1RIs7uHyIpgV96G" %}
[Guide: Configuring Custom Variables for Cluster Job Submission Scripts](/setup-configuration-and-management/software-system-guides/guide-configuring-custom-variables-for-cluster-job-submission-scripts)
{% endcontent-ref %}

## SSD Caching

{% content-ref url="/pages/-MNeFvdldyAikdOWQ4K3" %}
[Guide: SSD Particle Caching in CryoSPARC](/setup-configuration-and-management/software-system-guides/tutorial-ssd-particle-caching-in-cryosparc)
{% endcontent-ref %}

## Data Management, Import and Export

{% content-ref url="/pages/F3KBgDxkuaoVRFwpV0KW" %}
[Guide: Data Management in CryoSPARC (v4.0+)](/setup-configuration-and-management/software-system-guides/guide-data-management-in-cryosparc-v4.0)
{% endcontent-ref %}

{% content-ref url="/pages/Sgc4uptcwUJhNnXMk6v9" %}
[Guide: Data Cleanup (v4.3+)](/setup-configuration-and-management/software-system-guides/guide-data-cleanup-v4.3)
{% endcontent-ref %}

{% content-ref url="/pages/1LzAjuvhV2nWupJX9Gmd" %}
[Guide: Reduce Database Size (v4.3+)](/setup-configuration-and-management/software-system-guides/guide-reduce-database-size-v4.3)
{% endcontent-ref %}

{% content-ref url="/pages/-MNexr5lrR2HBCuP4qWS" %}
[Guide: CryoSPARC Live Session Data Management (≤v4.7)](/setup-configuration-and-management/software-system-guides/cryosparc-live-session-data-management-4.7)
{% endcontent-ref %}

{% content-ref url="/pages/kkC48tlWZTekJEslJFbr" %}
[Guide: Instance Recovery (v5.0+)](/setup-configuration-and-management/software-system-guides/guide-instance-recovery-v5.0)
{% endcontent-ref %}

{% content-ref url="/pages/-MNeMxOeXMsIWu8jTuYb" %}
[Guide: Migrating your CryoSPARC Instance](/setup-configuration-and-management/software-system-guides/tutorial-migrating-your-cryosparc-instance)
{% endcontent-ref %}


# Guide: Updating to CryoSPARC v5

CryoSPARC v5 is backwards compatible with v4. The update process includes new validation steps that may take some time, up to one hour for larger instances.

{% hint style="info" %}
As of May 27, 2026, the stable version of CryoSPARC v5 is available as v5.0.6.
{% endhint %}

## Updating to CryoSPARC v5 from v4

In v5, CryoSPARC’s underlying software system has been completely redesigned for stability, scalability, and to enable future developments. The new system adds strong validation and consistency guarantees for CryoSPARC database contents, minimizing the likelihood of errors and issues.

**CryoSPARC v5 is backwards compatible with v4 versions**, meaning:

* A v4.0+ instance can be upgraded to v5, and also downgraded back to v4.4+ if needed
* Projects detached from v4 instances can be attached to v5 instances, and vice versa

Updating to CryoSPARC v5 will perform a validation and upgrade of all database contents and will take some time, up to one hour for large instances. The user interface will not be available during the update process.

In the sections below you will find:

* [Compatibility requirements including OS and NVIDIA Driver versions for CryoSPARC v5](#compatibility-requirements)
* [A walkthrough of the update process that includes validation of database contents](#walkthrough-of-update-process)
* [Steps and commands to update from CryoSPARC v4 to CryoSPARC v5](#steps-and-commands-to-update-from-v4-to-v5)
* [Steps and commands to downgrade from CryoSPARC v5 to CryoSPARC v4](#downgrade-from-cryosparc-v5-to-v4)

### Compatibility Requirements

* **CryoSPARC Versions:** You must be running CryoSPARC v4.0+ to update to v5. Updating from CryoSPARC v3 or below is not supported. For older instances, please update to a v4 version first, then update to v5. Once v5 is installed, you can downgrade to CryoSPARC versions v4.4+; If you need to downgrade further back than v4.4, downgrade to v4.7 first, then downgrade to the older version.
* **Operating System:** For CryoSPARC v5, [the operating system must support GLIBC 2.28 or greater](https://guide.cryosparc.com/setup-configuration-and-management/hardware-and-system-requirements#operating-system). Therefore, the oldest compatible operating systems are Rocky/RHEL 8 and Ubuntu 20.04. For Ubuntu, version 22.04 or 24.04 is recommended. [*As of June 2026, CryoSPARC versions up to 5.0.6 are incompatible with the version 7 kernel that is the default for Ubuntu 26.04.*](#user-content-fn-1)[^1]
* **NVIDIA Driver:** [CryoSPARC v5 requires NVIDIA driver version 570.26 or newer](https://guide.cryosparc.com/setup-configuration-and-management/cryosparc-installation-prerequisites). Note that NVIDIA Blackwell devices are only compatible with the open driver. CryoSPARC v5 uses CUDA 12.8 which drops support for NVIDIA GPUs with compute capability 3.5 (Kepler). Only GPUs with compute capability 5.0 (Maxwell) to 12.0 (Blackwell) are supported.
* **cryosparc-tools:** A new, backwards compatible version of [cryosparc-tools](https://tools.cryosparc.com/intro.html) for scripting with v5 is available; see [details here](https://github.com/cryoem-uoft/cryosparc-tools/blob/main/CHANGELOG.md). Scripts written with previous versions will continue to function as before.
* **CryoSPARC CLI:** CryoSPARC v5 introduces a [new improved command line interface](https://guide.cryosparc.com/setup-configuration-and-management/management-and-monitoring-v5.0+/cryosparcm-cli-reference-v5.0+) that is not compatible with v4 `cli` commands. Scripts that used v4 `cli` commands **will need to be updated**, including for [managing CryoSPARC Live sessions](https://guide.cryosparc.com/live/managing-a-cryosparc-live-session-from-the-cli-v5.0+).

### Walkthrough of Update Process

CryoSPARC v5 includes new database validation checks that ensure consistency. For this reason, when updating from v4 to v5, a **database upgrade step** will be automatically performed during the update process. This step confirms the validity of all existing documents in the CryoSPARC database (e.g. users, projects, jobs) and potentially could **take up to an hour** for large instances. The database upgrade step runs in the command line as part of the update command; the **CryoSPARC user interface will not be available while it is in progress**. In rare cases, the database upgrade step may find some data in the database that is invalid (for example, if a user had manually modified the database in an inconsistent way in the past).

The database upgrade will proceed in **two phases**:

1. **Dry run validation phase** (no modifications are made to the database during this phase):
   * Database contents are validated, including users, scheduler targets, projects, workspaces, jobs, sessions, and exposures. The validation process will keep track of any invalid data that it finds that it can fix automatically. These pending fixes will be written to a validation JSON file, e.g.: `cryosparc_master/run/validation_results_2025_07_30_23h50.json`.
   * If the validation phase finds invalid database contents within a project (i.e., jobs, workspaces, sessions) then the user is given the option to detach the corresponding invalid project as part of the next phase. No data on disk will be deleted, and detached projects can be attached to another CryoSPARC v4 instance. If the user does not wish to detach invalid projects, the upgrade is aborted: [downgrade to v4 using the instructions below.](#downgrade-from-cryosparc-v5-to-v4)
   * If the validation phase finds other invalid database content that it cannot fix, the upgrade is aborted: [downgrade to v4 using the instructions below](#downgrade-from-cryosparc-v5-to-v4). The update process will ask if it is okay to upload the validation JSON file to Structura to helps us identify and fix any issues.
2. **Upgrade write phase** (changes are written to the database and disk):
   * Upgrade proceeds to make the changes that were validated in the previous step. Change results are written to an upgrade JSON file, e.g.: `cryosparc_master/run/upgrade_results_2025_07_30_23h50.json`.

Once the update is complete, CryoSPARC v5 starts and is ready for use.

### Steps and commands to update from v4 to v5

{% hint style="danger" %}
Do not attempt installing a new v5 instance with an existing CryoSPARC database that was previously in-use by v4. You must run `cryosparcm update` to complete database validation. Skipping the update step may result in data loss.
{% endhint %}

1. Ensure all [compatibility requirements](#compatibility-requirements) are met.
2. Follow the standard [Before you update or downgrade](/setup-configuration-and-management/software-updates#before-you-update-or-downgrade) steps to ensure you have sufficient disk space, have killed running jobs, completely shut down CryoSPARC, and created a backup of the CryoSPARC database.
3. Run `cryosparcm update`
   1. This command will update to the latest CryoSPARC v5 release. It will download the new version, update dependencies, then proceed with the database upgrade step described above before starting the CryoSPARC instance.

## Downgrade from CryoSPARC v5 to v4

Run `cryosparcm update --version v4.7.1` (or replace the version with another version ≥ `v4.4.0`, if desired). The instance can be started and run normally after this.

{% hint style="info" %}
Note that it is not possible to downgrade from v5.0 directly to any version older than v4.4. If you need to downgrade further back than v4.4, downgrade to v4.7 first then downgrade to the older version.
{% endhint %}

[^1]:


# Guide: Updating to CryoSPARC v4

Installing or updating to CryoSPARC v4 is similar to previous versions of CryoSPARC, but downgrading is not possible past v3.4.0.

### Upgrading from v3 to v4[​](https://beta.cryosparc.com/docs/installing-upgrading#upgrading-from-v3-to-v4) <a href="#upgrading-from-v3-to-v4" id="upgrading-from-v3-to-v4"></a>

All CryoSPARC projects, jobs and Live Sessions created in CryoSPARC v3 are forwards compatible with v4. You can upgrade an existing CryoSPARC v3 instance to CryoSPARC v4 using the same process outlined here [Software Updates and Patches](/setup-configuration-and-management/software-updates#installing-automatic-updates)

{% hint style="danger" %}
Because CryoSPARC v4.0+ relies on a newer version of MongoDB, after upgrading it will not be possible to downgrade to a CryoSPARC version below v3.4.0.
{% endhint %}

1. ⚠️ [Make a backup of your database](https://guide.cryosparc.com/setup-configuration-and-management/management-and-monitoring/cryosparcm#cryosparcm-backup)
2. `cryosparcm update`

{% hint style="info" %}
This command will update to the latest version of CryoSPARC.

To update to a specific version, add e.g.,`--version=v4.0.0`
{% endhint %}

### Downgrading from v4[​](https://beta.cryosparc.com/docs/installing-upgrading#downgrading) <a href="#downgrading" id="downgrading"></a>

CryoSPARC v4.0+ relies on a newer version of MongoDB, v3.6, than CryoSPARC v3.3, which relies on MongoDB 3.4. Therefore, after upgrading to CryoSPARC v4.0, it will no longer be possible to downgrade to a version of CryoSPARC that relies on a version of MongoDB older than v3.6.

CryoSPARC v3.4.0 was created to allow downgrades from a v4 version of CryoSPARC to a v3 version, where necessary. CryoSPARC v3.4.0 is a carbon copy of CryoSPARC v3.3.2+220824, the most recent pre-v4.0 release of CryoSPARC, and includes support for MongoDB v3.6. It is possible to update to CryoSPARC v4.0 from v3.4.

1. ⚠️ [Make a backup of your database](https://guide.cryosparc.com/setup-configuration-and-management/management-and-monitoring/cryosparcm#cryosparcm-backup)
2. [Follow downgrade instructions from the guide](https://guide.cryosparc.com/setup-configuration-and-management/software-updates#update-or-roll-back-downgrade-to-a-specific-version) and specify: `--version=v3.4.0`

### Application & Port Changes[​](https://beta.cryosparc.com/docs/installing-upgrading#port-changes) <a href="#port-changes" id="port-changes"></a>

The new web application replaces the legacy web application at the main (base) port of your CryoSPARC instance. In v4+, the legacy web application is not started by default during `cryosparcm start`, but can be turned on by running `cryosparcm start app_legacy`.

You can still access the legacy web application via these ports:

| Application        | Port Number                  |
| ------------------ | ---------------------------- |
| New application    | `BASE_PORT` (e.g 39000)      |
| Legacy application | `BASE_PORT + 7` (e.g. 39007) |

[<br>](https://beta.cryosparc.com/docs/intro)


# Guide: Installation Testing with cryosparcm test

This guide covers how to use cryosparcm test to verify your CryoSPARC installation is working properly.

{% hint style="warning" %}
The information in this section applies to CryoSPARC v4.0+.
{% endhint %}

## Overview

After installing CryoSPARC [using the instructions here](https://guide.cryosparc.com/setup-configuration-and-management/how-to-download-install-and-configure), you can verify your instance is correctly installed by using `cryosparcm test install` and `cryosparcm test workers` via the command line. Running these functions will perform several tests that ensure users can seamlessly launch jobs and process data in CryoSPARC.

## `cryosparcm test install`

This function tests the core components of CryoSPARC (HTTP connections, licensing, workers, etc.) that are required to start running jobs and provides information on the status of the CryoSPARC instance (e.g., which version is running, whether a patch is available, etc.).

To run this function, log into a shell on the master node as the user that owns the CryoSPARC instance.

Run `cryosparcm test -h` for full usage instructions.

### Example Output

```
cryosparcuser@uoft ~/ $ cryosparcm test i
✓ Running as CryoSPARC owner cryosparcuser
✓ Running on master node uoft
✓ CryoSPARC is running
✓ CRYOSPARC_LICENSE_ID environment variable is set
✓ Insecure mode is disabled
✓ License server set to "https://get.cryosparc.com"
✓ Connection to license server succeeded
✓ License server returned success status code 200
✓ License server returned valid JSON response
✓ License exists and is valid
✓ CryoSPARC is running v5.0.0
✓ Develop version - no updates available
✓ Admin user exists
✓ GPU worker connected
```

### Test Checklist

Running `cryosparcm test install` will test or check the following components:

1. Test if the `cryosparcm test install` command is running as the user who owns the CryoSPARC instance.
2. Test if the `cryosparcm test install` command is running on the machine that runs the CryoSPARC master instance.
3. Check if the CryoSPARC instance is turned on.
   * If this test fails, turn on CryoSPARC by running `cryosparcm start` and run the command again.
4. Test if an HTTP connection can be successfully created to the `api` (`CRYOSPARC_BASE_PORT`+2) server.
   * If this test fails, ensure a firewall isn’t blocking access to the ten consecutive ports from `CRYOSPARC_BASE_PORT` (default 61000, e.g., 61000-61010). For more information, see [Open TCP Ports](https://guide.cryosparc.com/setup-configuration-and-management/cryosparc-installation-prerequisites#4.-open-tcp-ports) in the Guide.
5. Check if the environment variable `CRYOSPARC_LICENSE_ID` is set.
6. Test if the CryoSPARC License ID is in the correct format.
   1. If this test fails, ensure the CryoSPARC License ID found in `cryosparc_master/config.sh` is set to the correct license ID.
7. Check if insecure request mode is enabled or disabled.
   1. This option is controlled by the `CRYOSPARC_INSECURE` environment variable found in `cryosparc_master/config.sh`.
   2. Enabling this option ignores SSL certificate errors when connecting to HTTPS endpoints. This is useful if you are behind an enterprise network using SSL injection.
8. Check if the URL to the license server is valid.
   1. The URL can be overridden by the `CRYOSPARC_LICENSE_SERVER_ADDR` environment variable found in `cryosparc_master/config.sh`.
   2. The default URL is <https://get.cryosparc.com>
9. Check if the CryoSPARC License ID being used is active.
   1. If this test fails, either the CryoSPARC instance wasn’t able to connect to the licensing server, the license isn’t active, or there was a network partition causing data corruption (in which case, trying the command again in a few minutes may fix the issue).
   2. If the instance is having trouble connecting to the licensing server, see [License Server Troubleshooting](https://guide.cryosparc.com/setup-configuration-and-management/troubleshooting#license-error-or-license-not-found) in the Guide. Additionally, if your network is behind an HTTP proxy, see [Custom SSL Certificate Authority Bundle](https://guide.cryosparc.com/setup-configuration-and-management/how-to-download-install-and-configure/accessing-cryosparc#appendix-d-custom-ssl-certificate-authority-bundle) in the Guide.
   3. If the license being used is no longer active, request a new CryoSPARC License ID. See [Obtaining A License ID](https://guide.cryosparc.com/setup-configuration-and-management/how-to-download-install-and-configure/obtaining-a-license-id) in the Guide.
10. Check the current running version of the CryoSPARC instance.
    1. See the [CryoSPARC Changelog](https://cryosparc.com/updates).
11. Check if there is an update available for CryoSPARC.
    1. To update CryoSPARC, run `cryosparcm update`. For more information, see [Software Updates and Patches](https://guide.cryosparc.com/setup-configuration-and-management/software-updates) in the Guide.
12. Check if there is a patch update available for CryoSPARC.
    1. To patch CryoSPARC, run `cryosparcm patch`. For more information, see [Apply Patches](https://guide.cryosparc.com/setup-configuration-and-management/software-updates#apply-patches) in the Guide.
13. Check if a worker is connected with at least one GPU.
    1. To add a GPU worker to CryoSPARC, see [Connecting A Worker Node](https://guide.cryosparc.com/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc#connecting-a-worker-node) in the Guide.

## `cryosparcm test workers`

This function tests workers connected to CryoSPARC to ensure they can correctly run CryoSPARC jobs by testing if the worker can launch jobs, cache particles to an SSD (if an SSD is configured), and utilize the GPU correctly. This test can be run via the command line, or directly in the CryoSPARC user interface. Three new jobs have been added to CryoSPARC that can be run at any time on the lane you’d like to test.

<figure><img src="/files/KsjCba0CAaQM6oGqBVD2" alt=""><figcaption><p>You can find the jobs used for worker tests in the "Instance Testing Utilities" section of the job builder.</p></figcaption></figure>

### Usage

Run `cryosparcm test --help` for full usage instructions.

The tests require a project to be run inside. If there are no projects in the instance, create one before running this function.

To run all tests on all workers:

* run `cryosparcm test workers <project_uid> --test all`
* (e.g., `cryosparcm test workers P1 --test all`)

To run only the GPU test on all workers:

* run `cryosparcm test workers <project_uid> --test gpu`
* (e.g., `cryosparcm test workers P1 --test gpu`).

To run only the GPU test on a single worker:

* run `cryosparcm test workers <project_uid> --test gpu --target <workstation_hostname>`
* (e.g., `cryosparcm test workers P1 --test gpu --target cryoem1.uoft.ca`)

To run only the GPU test (with Tensorflow and PyTorch) on a single worker:

* run `cryosparcm test workers <project_uid> --test gpu --test-tensorflow --test-pytorch --target <workstation_hostname>`
* (e.g., `cryosparcm test workers P1 --test gpu --test-tensorflow --test-pytorch --target cryoem1.uoft.ca`)

To run only the GPU test on two workers:

* run `cryosparcm test workers <project_uid> --test gpu --target <workstation1_hostname> --target <workstation2_hostname>`
* (e.g., `cryosparcm test workers P1 --test gpu --target cryoem1.uoft.ca --target cryoem2.uoft.ca`)

### Example Output

*Some text removed for readability.*

```
cryossparcuser@uoft ~/ $ cryosparcm test workers P1
Using project P1
Running worker tests...
Worker test results
cryoem3
  ✓ LAUNCH
  ✓ SSD
  ✓ GPU
cryoem2
  ✓ LAUNCH
  ✓ SSD
  ✓ GPU
cryoem5
  ✓ LAUNCH
  ✕ SSD
    Error: [Errno 13] Permission denied: '/scratch'
    See P1 J1211 for more information
  ⚠ GPU
    No GPU available
cryoem6
  ✕ LAUNCH
    Error: 
    See P1 J1203 for more information
  ⚠ SSD
    Did not run: Launch test failed
  ⚠ GPU
    Did not run: Launch test failed
cryoem1
  ✓ LAUNCH
  ✓ SSD
  ✓ GPU
    ⚠ RTX A6000 @ 00000000:03:00.0: Persistence Mode is Disabled. 
      Enable Persistence mode by running `nvidia-smi -pm 1` as root to persist 
      the NVIDIA driver, reducing GPU load times.
    ⚠ RTX A6000 @ 00000000:03:00.0: GPU Software Power Cap is Active
    ⚠ RTX A6000 @ 00000000:21:00.0: Persistence Mode is Disabled. 
      Enable Persistence mode by running `nvidia-smi -pm 1` as root to persist 
      the NVIDIA driver, reducing GPU load times.
    ⚠ RTX A6000 @ 00000000:21:00.0: GPU Software Power Cap is Active
    ⚠ GeForce RTX 3090 @ 00000000:4C:00.0: Persistence Mode is Disabled. 
      Enable Persistence mode by running `nvidia-smi -pm 1` as root to persist 
      the NVIDIA driver, reducing GPU load times.
cryoem7
  ✓ LAUNCH
  ✓ SSD
  ✓ GPU
cryoem9
  ✓ LAUNCH
  ✓ SSD
  ✕ GPU
    Error: Tensorflow detected 0 of 7 GPUs.
    See P1 J1222 for more information
cryoem10
  ✓ LAUNCH
  ✓ SSD
  ✓ GPU
```

When the worker test is run, a new workspace inside the specified project will be created to contain all test jobs. The workspace will be named with the date and time (UTC) of execution.

<figure><img src="/files/c39okPINvAXSnPzEW8mB" alt=""><figcaption><p>Workspace card of an instance testing run.</p></figcaption></figure>

{% hint style="info" %}
If a test job fails, check the job's Event Log and [stdout log (joblog)](https://guide.cryosparc.com/setup-configuration-and-management/management-and-monitoring/cryosparcm?q=queuing+#cryosparcm-joblog-px-jxx) for more details.
{% endhint %}

### Launch Test

The ability to launch jobs will be tested first. This will indicate if the worker is accessible and can correctly run CryoSPARC jobs. If this test fails, it most likely indicates a connection issue between the master and the worker. For more information, see [Cannot Queue or Run Job](https://guide.cryosparc.com/setup-configuration-and-management/troubleshooting#cannot-queue-or-run-job) in the Guide.

Note that if a launch test fails on a worker, the SSD and GPU tests will not run:

```
cryoem6
  ✕ LAUNCH
    Error: ssh: connect to host cryoem6 port 22: No route to host
    See P1 J1203 for more information
  ⚠ SSD
    Did not run: Launch test failed
  ⚠ GPU
    Did not run: Launch test failed
```

### SSD Test

If an SSD is configured for a worker, the SSD test will confirm that particle caching is working properly. The test creates five different particle stacks of shape `(500, 512, 512)` in the project directory, and tries to cache them to the SSD.

```
Testing SSD

Generating a 500 particle stack with shape (512, 512).

Writing particle stack 1/5... Done in 1.517s
Writing particle stack 2/5... Done in 1.329s.
Writing particle stack 3/5... Done in 1.290s.
Writing particle stack 4/5... Done in 1.221s.
Writing particle stack 5/5... Done in 1.219s.

Loading a ParticleStack with 5 items...
 SSD cache : cache successfully synced in_use
 SSD cache : cache successfully synced, found 233127.92MB of files on SSD.
 SSD cache : cache successfully requested to check 5 files.
 SSD cache : cache requires 2500.00MB more on the SSD for files to be downloaded.
 SSD cache : cache has enough available space.

 Transferring J33/data/simulated_particles_4.mrc (500 MB) (5/5)
  Complete      :         2500 MB (100.00%)
  Total         :         2500 MB
  Current Speed :    1133.22 MB/s
  Average Speed :    1089.09 MB/s
  ETA           :      0h  0m  0s

 SSD cache : complete, all requested files are available on SSD.
Done.

Cleaning up testing data...
SSD Test completed successfully.
```

If an SSD Test fails for any reason, the reason will be summarized in the test results:

```
cryoem5
  ✓ LAUNCH
  ✕ SSD
    Error: [Errno 13] Permission denied: '/scratch'
    See P1 J1211 for more information
  ⚠ GPU
    No GPU available
```

For more information on configuring and troubleshooting an SSD cache for a worker, see [SSD Particle Caching in CryoSPARC](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/tutorial-ssd-particle-caching-in-cryosparc) in the Guide.

### GPU Test

The GPU test will collect information about all the GPUs on the worker and test if the worker can compile and run GPU code.

The following information is collected about each GPU via `nvidia-smi`:

* `driver_version`: GPU driver version
  * keeping the driver up to date ensures stability
  * [NVIDIA Driver Downloads](https://www.nvidia.com/Download/index.aspx)
* `persistence_mode`: GPU driver persistence
  * [NVIDIA Docs: Driver Persistence](https://docs.nvidia.com/deploy/driver-persistence/index.html)
  * enabling this reduces GPU driver load times
  * enable this by running `nvidia-smi --pm 1` as root
* `power_limit`: GPU power limit (TDP)
  * information only
* `sw_power_limit`: software power limiter
  * if “Active”, this might indicate the power supply unit (PSU) on the workstation isn’t able to support the power draw from the GPU, or if a power supply cable is faulty or not properly connected to the GPU
  * if “Active”, this might indicate the GPU temperature is too high
* `hw_power_limit`: hardware power limiter
  * if “Active”, this might indicate the power supply unit (PSU) on the workstation isn’t able to support the power draw from the GPU
  * if “Active”, this might indicate the GPU temperature is too high
* `compute_mode`: current compute mode (Default, Exclusive Process, etc.)
  * the “default” compute mode allows users to launch multiple GPU jobs onto the same GPU via the Queue modal in the UI. See [Queuing Directly To A GPU](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/tutorial-queuing-directly-to-a-gpu?q=queuing+) in the Guide.
  * the “exclusive process” compute mode prevents a process from obtaining a context from a GPU while another process already has one, useful in anonymous multi-user scenarios
  * set the compute mode of the GPU by running `nvidia-smi -c compute_mode -i target_gpu_id` where `compute_mode` is one of:
    * 0/Default, 1/Exclusive Thread, 2/Prohibited, 3/Exclusive Process
* `max_pcie_link_gen`: maximum PCIe link generation (e.g., PCIe 3 or PCIe 4)
  * information only
* `current_pcie_link_gen`: current PCIe link generation
  * information only
  * this may be equal to or lower than the `max_pcie_link_gen`, as the GPU automatically switches to a higher link under load
* `temperature`: current temperature
  * information only
* `gpu_utilization`: current utilization
  * information only
* `memory_utilization`: current memory utilization
  * information only

Example data:

```
Obtaining GPU info via `nvidia-smi`...

NVIDIA GeForce RTX 3090 @ 00000000:01:00.0
    driver_version                :510.68.02
    persistence_mode              :Enabled
    power_limit                   :350.00
    sw_power_limit                :Not Active
    hw_power_limit                :Not Active
    compute_mode                  :Default
    max_pcie_link_gen             :4
    current_pcie_link_gen         :1
    temperature                   :25
    gpu_utilization               :0
    memory_utilization            :0

NVIDIA A100-PCIE-40GB @ 00000000:61:00.0
    driver_version                :510.68.02
    persistence_mode              :Enabled
    power_limit                   :250.00
    sw_power_limit                :Not Active
    hw_power_limit                :Not Active
    compute_mode                  :Default
    max_pcie_link_gen             :4
    current_pcie_link_gen         :4
    temperature                   :33
    gpu_utilization               :0
    memory_utilization            :0

Starting PyCuda GPU test on: NVIDIA A100-PCIE-40GB @ 0000:61:00.0
    PyCuda was compiled with CUDA: (11, 2, 0)
Finished PyCuda GPU test in 0.026s

Testing Tensorflow...
    Tensorflow found 4 GPUs.
Tensorflow test completed in 3.385s
```

Finally, PyCUDA (and optionally Tensorflow and PyTorch) will be tested to ensure they are working properly. If the either of these tests fail, the error will be summarized in the test results. For more information, check the failed job’s Event Log and [stdout log (joblog)](https://guide.cryosparc.com/setup-configuration-and-management/management-and-monitoring/cryosparcm?q=queuing+#cryosparcm-joblog-px-jxx).

```
cryoem9
  ✓ LAUNCH
  ✓ SSD
  ✕ GPU
    Error: Tensorflow detected 0 of 7 GPUs.
    See P1 J1222 for more information
```

#### Testing Tensorflow and PyTorch

By default, Tensorflow and PyTorch capabilities are not tested during the GPU test. To enable these tests, specify `--test-tensorflow` and/or `--test-pytorch` when starting the worker test. For example:

`cryosparcm test workers P12 --test-tensorflow --target cryoem9.structura.dev`

{% hint style="info" %}
The PyTorch test will fail if the 3D Flex Refine dependencies were not installed using `cryosparcw install-3dflex` introduced in CryoSPARC v4.1.0. For more information, see \<Link to 3D Flex Refine: Installing Dependencies>
{% endhint %}

If Tensorflow or PyTorch was not able to detect all GPUs on your system, the job will fail, and the error message will appear in the job's stdout log (found in the 'Metadata' tab of the Job Dialog).


# Guide: Verify CryoSPARC Installation with the Extensive Validation Job (v4.3+)

{% hint style="info" %}
The "Extensive Workflow" job has been renamed to "Extensive Validation" in CryoSPARC v4.3.0+. For the version of this guide applicable to CryoSPARC versions ≤v4.2, please see: [Extensive Workflow](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/tutorial-verify-cryosparc-installation-with-the-extensive-workflow-sysadmin-guide-1)
{% endhint %}

{% hint style="success" %}
In CryoSPARC v5.0+, many job types have been added to the Extensive Validation Job's test set, and in advanced mode (see below), nearly all job types that exist are launched, providing for a comprehensive test of CryoSPARC.
{% endhint %}

## Introduction

CryoSPARC's "Extensive Validation" job orchestrates a full 3D target reconstruction for two datasets:

* T20S Proteasome ([EMPIAR-10025](http://pdbe.org/EMPIAR-10025)) from a small subset of movies (\~8GB)
* Tobacco Mosaic Virus ([EMPIAR-10305](https://www.ebi.ac.uk/empiar/EMPIAR-10305/))

CryoSPARC's engineering team uses this job to automatically test and benchmark CryoSPARC between releases.

System Administrators may use the Extensive Validation job to verify that CryoSPARC is correctly configured following a fresh installation or an update.

![Table view of a workspace with an Extensive Validation run](/files/k26yaPzYSrcvMPTUU5O0)

### Benchmarking vs. Testing

The Extensive Validation job has two modes: Benchmark mode and Testing mode.

<figure><img src="/files/Eu7lgVvdwpNbIhWYZecp" alt=""><figcaption><p>Switch the Extensive Workflow job from "Testing" mode to "Benchmark" mode using the drop-down menu.</p></figcaption></figure>

In Benchmark mode, jobs with pre-defined parameters run sequentially on the worker node in order. Each job accesses available system resources independently of other jobs to collect accurate runtime statistics. **Benchmark mode is useful for evaluating the overall performance of a worker node.**

In Testing mode, jobs run in parallel when possible. Multiple parameter combinations of each job are dispatched. **Testing mode is useful for ensuring that a CryoSPARC instance is installed correctly.**

Both Benchmark and Testing modes verify the following system requirements:

* CryoSPARC system and license installation
* Worker/Cluster configuration
* GPU and CUDA driver installation
* SSD caching

### Datasets Available

{% hint style="info" %}
CryoSPARC downloads the selected dataset into the project directory when the Extensive Validation job first runs in the current project.
{% endhint %}

#### [EMPIAR-10025](https://www.ebi.ac.uk/empiar/EMPIAR-10025/)

* **Number of movies:** 20
* **Frames per movie:** 38
* **Movie size:** 7420 × 7676 (K2 Super Resolution)
* **Pixel size:** 0.86 Å
* **Particles processed:** 10,000
* **Particle box size (pixels):** 448

<details>

<summary>Jobs run in Benchmark Mode (up to CryoSPARC v4.7)</summary>

1. Import Movies
2. Patch Motion Correction
3. Patch CTF Estimation
4. Manually Curate Exposures
5. Blob Picker
6. Inspect Particle Picks
7. Extract From Micrographs (CPU)
8. 2D Classification
9. Select 2D
10. Template Picker
11. Inspect Particle Picks
12. Extract From Micrographs (CPU)
13. 2D Classification (50 Class)
14. 2D Classification (100 Class) (All Job Types Enabled)
15. 2D Classification (200 Class) (All Job Types Enabled)
16. Select 2D
17. Particle Sets Tools
18. Ab-Initio Reconstruction
19. Ab-Initio Reconstruction (3 Class) (All Job Types Enabled)
20. Homogeneous Refinement
21. Non-Uniform Refinement
22. Homogeneous Refinement Legacy (All Job Types Enabled)
23. Non-Uniform Refinement Legacy (All Job Types Enabled)
24. 3D Classification
25. 3D Variability (3 Mode)
26. 3D Variability (6 Mode) (All Job Types Enabled)

</details>

<details>

<summary>Jobs run in Testing Mode (up to CryoSPARC v4.7)</summary>

1. Import Movies
2. Patch Motion Correction
3. Full-Frame Motion Correction
4. Patch CTF Estimation
5. CTFFIND4
6. Curate Exposures (Stream A)
7. Curate Exposures (Stream B)
8. Blob Picker (Stream A)
9. Blob Picker (Stream B)
10. Inspect Picks (Stream A)
11. Inspect Picks (Stream B)
12. Extract From Micrographs (Stream A)
13. Local Motion Correction (Stream B)
14. 2D Classification (Stream A) # all jobs past this point are in Stream A
15. Select 2D
16. Template Picker
17. Inspect Picks
18. Extract From Micrographs
19. 2D Classification (50 Class)
20. 2D Classification (100 Class)
21. 2D Classification (200 Class)
22. Select 2D
23. Particle Sets Tools
24. Ab-Initio Reconstruction
25. Ab-Initio Reconstruction (3 Class)
26. Homogeneous Refinement
27. Non-Uniform Refinement
28. Homogeneous Refinement (Legacy)
29. Non-Uniform Refinement (Legacy)
30. Heterogeneous Refinement (3 Class)
31. Heterogeneous Refinement (6 Class)
32. 3D Classification (Simple mode)
33. 3D Classification (PCA mode)
34. 3D Variability (3 mode)
35. 3D Variability (6 mode)
36. Sharpening Tools
37. Validation (FSC)
38. Global CTF Refinement
39. Local CTF Refinement
40. 3D Variability Display

</details>

<details>

<summary>Jobs run in Benchmark Mode (CryoSPARC v5.0+)</summary>

1. Import Movies
2. Patch Motion Correction
3. Patch CTF Estimation
4. Manually Curate Exposures
5. Blob Picker
6. Inspect Particle Picks
7. Extract From Micrographs (CPU)
8. Micrograph Junk Detector (v5.0+)
9. 2D Classification
10. Downsample Particles (v5.0+)
11. Select 2D
12. Template Picker
13. Inspect Particle Picks
14. Extract From Micrographs (CPU)
15. 2D Classification (50 Class)
16. 2D Classification (100 Class) (All Job Types Enabled)
17. 2D Classification (200 Class) (All Job Types Enabled)
18. Select 2D
19. Particle Sets Tools
20. Ab-Initio Reconstruction
21. Ab-Initio Reconstruction (3 Class) (All Job Types Enabled)
22. Heterogeneous Refinement (v5.0+)
23. Homogeneous Refinement
24. Non-Uniform Refinement
25. Subset Particles by Statistic
26. Volume Tools
27. Orientation Diagnostics
28. 3D Variability (3 Mode)
29. 3D Variability (6 Mode) (All Job Types Enabled)
30. Local Refinement

</details>

<details>

<summary>Jobs run in Testing Mode (CryoSPARC v5.0+)</summary>

1. Import Movies
2. Import Micrographs
3. Import 3D Volumes
4. Import Particle Stack
5. Import Templates
6. Patch Motion Correction
7. Full-Frame Motion Correction
8. Import Beam Shift
9. Patch CTF Estimation
10. CTFFIND4
11. Check For Corrupt Micrographs
12. Generate Micrograph Thumbnails
13. Curate Exposures (Stream A)
14. Curate Exposures (Stream B)
15. Exposure Tools
16. Exposure Sets Tool (Split)
17. Exposure Sets Tool (Intersect)
18. Blob Picker (Stream A)
19. Blob Picker (Stream B)
20. Inspect Picks (Stream A)
21. Remove Duplicate Particles
22. Blob Picker Tuner
23. Inspect Picks (Stream B)
24. Reassign Particles to Micrographs
25. Patch CTF Extraction
26. Local Motion Correction (Stream B)
27. Local Motion Correction (Multi) (Stream B)
28. Extract From Micrographs (CPU) (Stream A)
29. Extract From Micrographs (GPU)
30. Micrograph Junk Detector
31. Restack Particles
32. 2D Classification (Stream A) # all jobs past this point are in Stream A
33. Cache Particles on SSD
34. Check For Corrupt Particles
35. Class Probability Filter
36. Downsample Particles
37. Select 2D
38. Average Power Spectra
39. Template Picker
40. Inspect Picks
41. Extract From Micrographs
42. Exposure Group Utilities (Combine Particles)
43. Exposure Group Utilities (Split Particles)
44. Exposure Group Utilities (Combine Exposures)
45. Exposure Group Utilities (Split Exposures)
46. 2D Classification (50 Class)
47. 2D Classification (100 Class)
48. 2D Classification (200 Class)
49. Rebalance 2D Classes
50. Select 2D
51. Particle Sets Tools
52. Reconstruct 2D Classes
53. Ab-Initio Reconstruction
54. Ab-Initio Reconstruction (3 Class)
55. Split Volumes Group
56. Simulate Data
57. Reference Based Auto Select 3D
58. Volume Tools
59. Create Templates (Non-Helical)
60. Create Templates (Helical)
61. Heterogeneous Refinement
62. Homogeneous Refinement
63. Non-Uniform Refinement
64. Subset Particles by Statistic
65. Volume Alignment Tools
66. Volume Tools (Mask)
67. Reference Based Auto Select 2D (Sobel)
68. Reference Based Auto Select 2D (Cluster)
69. Reference Based Auto Select 2D (Thresholds)
70. Rebalance Orientations
71. Orientation Diagnostics
72. Heterogenous Reconstruction Only
73. Local Resolution Estimation
74. Symmetry Search Utility
75. Apply Helical Symmetry
76. Reference Based Motion Correction
77. 3D Flex Data Prep
78. Patch Motion to Local Motion
79. 3D Classification
80. Sharpening Tools
81. Validation (FSC)
82. Global CTF Refinement
83. 3D Variability (3 mode)
84. 3D Variability (6 mode)
85. Local Refinement
86. Recenter Trajectories
87. Apply Trajectories
88. Local Filtering
89. 3D Flex Mesh Prep
90. Align 3D Maps
91. Regroup 3D Classes
92. Particle Subtraction
93. 3D Flex Training
94. Local CTF Refinement
95. ResLog Analysis
96. 3D Variability Display
97. 3D Flex Generator
98. 3D Flex Reconstruction

</details>

#### [EMPIAR-10305](https://www.ebi.ac.uk/empiar/EMPIAR-10305/)

* **Number of movies:** 62
* **Frames per movie:** 20
* **Movie size:** 7420 × 7676 (K2 Super Resolution)
* **Pixel size:** 0.32 Å
* **Particles processed:** \~30,000
* **Particle box size (pixels):** 512

<details>

<summary>Jobs run in Benchmark Mode (up to CryoSPARC v4.7)</summary>

1. Import Movies
2. Patch Motion Correction
3. Patch CTF Estimation
4. Curate Exposures
5. Blob Picker
6. Inspect Picks
7. Extract From Micrographs
8. 2D Classification
9. Select 2D
10. Template Picker
11. Inspect Picks
12. Extract From Micrographs
13. 2D Classification
14. Select 2D
15. Particle Sets Tools
16. Ab-Initio Reconstruction
17. Homogeneous Refinement
18. Non-Uniform Refinement
19. 3D Classification
20. 3D Variability

</details>

<details>

<summary>Jobs run in Testing Mode (up to CryoSPARC v4.7)</summary>

1. Import Movies
2. Patch Motion Correction
3. Patch CTF Estimation
4. Filament Tracer
5. Inspect Picks
6. Extract From Micrographs
7. 2D Classification
8. Select 2D
9. Helical Refinement
10. Local CTF Refinement
11. Global CTF Refinement
12. Symmetry Expansion
13. Homogeneous Reconstruct Only

</details>

<details>

<summary>Jobs run in Benchmark Mode (CryoSPARC v5.0+)</summary>

1. Import Movies
2. Patch Motion Correction
3. Patch CTF Estimation
4. Filament Tracer
5. Inspect Particle Picks
6. Extract From Micrographs (CPU)
7. 2D Classification
8. Select 2D Classes
9. Helical Refinement
10. Local CTF Refinement
11. Global CTF Refinement
12. Symmetry Expansion
13. Homogeneous Reconstruction Only

</details>

<details>

<summary>Jobs run in Testing Mode (CryoSPARC v5.0+)</summary>

1. Import Movies
2. Patch Motion Correction
3. Patch CTF Estimation
4. Filament Tracer
5. Inspect Particle Picks
6. Extract From Micrographs (CPU)
7. 2D Classification
8. Select 2D Classes
9. Helical Refinement
10. Local CTF Refinement
11. Global CTF Refinement
12. Symmetry Expansion
13. Homogeneous Reconstruction Only

</details>

## Prerequisites

{% content-ref url="/pages/-M7DHIJsaWhZ95BFAmfp" %}
[How to Download, Install and Configure](/setup-configuration-and-management/how-to-download-install-and-configure)
{% endcontent-ref %}

## Creating and Running the Extensive Validation Job

1. Open the CryoSPARC web interface
2. In the dashboard, create a new Project from the navigation bar and create an initial workspace.

![](/files/6cQbGcdlyv6FyhfdYpNs)

Specify a descriptive title such as "Extensive Validation Testing" and directory for the project to store its data.

<figure><img src="/files/zSxsPu5mmdW6W75Vb0W4" alt=""><figcaption></figcaption></figure>

**Best practices:** Create a new workspace and run the Extensive Validation in that workspace each time CryoSPARC updates and restarts. Name each workspace with the latest installed version of CryoSPARC that the job runs on. For example, when testing CryoSPARC v4.3.0, name the workspace "v4.3.0 Benchmark and Validation"

4\. Select the Job Builder from the sidebar and select the **"Extensive Validation"** job (under the Validation category).

![](/files/DlF3WPH3nzplWw6fE0bD)

*5. (Optional)* If desired, change the job parameters.

6\. Select the node or cluster that the Extensive Validation jobs should run on.

![](/files/GmpOPDLtRFToqmbFH5zX)

![](/files/34wWzotnPo7iPCKCiUXt)

Queue the job from the Job Builder sidebar. Open the job's Event log (either click/tap the Job card header or select the Job card and press the `Space` key) to monitor its progress. The Validation job logs each spawned job as it is queued and logs how long it takes to complete.

![](/files/ZqXkgKgfcnHlcZXfy5pV)

Close the Job modal with the `×` button. This shows the workspace overview with CryoSPARC jobs spawned by the Extensive Validation job

Once all jobs successfully complete, the Extensive Validation job status changes to "Completed". This means the installation was successful. Users may now be notified to start or resume processing!

## Troubleshooting Failed Jobs

Extensive Validation will fails if any spawned job fails.

Scroll through the workspace to find jobs with the "Failed" status. Open the failed job's Event Log. Scroll to the bottom to see why the job failed.

Common failure reasons include

* [Cannot verify license](https://discuss.cryosparc.com/t/new-install-error-connecting-to-cryosparc-license-server/3825) or license key entered incorrectly
* [Incorrect filesystem permissions](https://discuss.cryosparc.com/t/file-r-w-permissions/3544)
* [CUDA not set up correctly](https://discuss.cryosparc.com/t/3d-variability-analysis-errors-v2-9-0/3110)
* [SSD cache full or not set up correctly](https://discuss.cryosparc.com/t/changing-ssd-no-run-directory/2916)
* [Worker registration issue](https://discuss.cryosparc.com/t/exit-status-255-error/2177)
* [Not enough GPU memory available](https://discuss.cryosparc.com/t/patchmotion-failure-2-13-2/3971)

Once the configuration issue is resolved, restart the Extensive Validation job: Either create a new workspace and job as already noted, or clear the existing Extensive Workflow job and re-queue.

## Additional Extensive Testing

For an even *more* extensive system test, the Extensive Validation job provides the parameter "Run Advanced Jobs"

![Turn on the "Run Advanced Jobs" radio button to run the full workflow.](/files/A8E8n5os4f2PWp5AirHR)

With "Run Advanced Jobs" enabled, additional validation jobs run in parallel. Use this to verify multi-GPU performance on a single node. Advanced jobs available for each dataset are listed in the "Datasets Available" section above.

## Expected Results

To compare the results of Extensive Validation runs, use "Benchmark" mode. This locks in the parameters and runs each job serially to ensure all system resources are available independently. The benchmark results may be viewed in the "Manage" panel, under the "Benchmarks" tab.

<figure><img src="/files/iiGdjfx5aoKewXAZkWJM" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/5pdcvX22UDUANSOtrjmM" alt=""><figcaption></figcaption></figure>

There are several reference benchmarks available for comparison with your CryoSPARC installation, including benchmarks completed on AWS EC2 instances. Select one or more benchmark rows and click "Compare" to compare benchmarks.

<figure><img src="/files/t03gEWdx3WoyFWEYFxBn" alt=""><figcaption></figcaption></figure>


# Guide: Verify CryoSPARC Installation with the Extensive Workflow (≤v4.2)

{% hint style="info" %}
For the version of this guide applicable to CryoSPARC versions v4.3.0+, please see: [Extensive Validation](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/tutorial-verify-cryosparc-installation-with-the-extensive-workflow-sysadmin-guide)
{% endhint %}

## Introduction

CryoSPARC provides a job called "Extensive Workflow for T20S", which performs a full 3D reconstruction of the T20S Proteasome ([EMPIAR-10025](http://pdbe.org/EMPIAR-10025)) from a small (\~8GB) subset of movies. The CryoSPARC engineering team uses this job to automatically test and benchmark CryoSPARC between releases.

System Administrators may use the extensive workflow job to verify that CryoSPARC is correctly configured following a fresh installation or an update.

![](/files/-MNhtI83mGjXUgYrnijl)

The Extensive Workflow covers the [full T20S tutorial](/guides-for-v3/cryo-em-data-processing-in-cryosparc-introductory-tutorial), including the following workflow jobs:

* Import Micrographs
* Motion Correction
* CTF Estimation
* Particle Picking and Extraction
* 2D Classification
* *Ab-initio* reconstruction
* Homogeneous Refinement

The following system requirements are verified:

* CryoSPARC system and license installation
* Worker/Cluster configuration
* GPU and CUDA driver installation
* SSD caching

The sample data has the following characteristics:

* **Number of images:** 20
* **Frames per image:** 38
* **Image size:** 7420 × 7676 (K2 Super Resolution)
* **Pixel size:** 0.66 Å

Once started, the workflow should take no more than an hour to complete.

## Prerequisites

{% content-ref url="/pages/-M7DHIJsaWhZ95BFAmfp" %}
[How to Download, Install and Configure](/setup-configuration-and-management/how-to-download-install-and-configure)
{% endcontent-ref %}

## Creating and Running the Extensive Workflow

1. Open the CryoSPARC web interface
2. In the dashboard, create a new Project from the navigation bar

![](/files/-MNhtOMDdmuUq0CAHwq6)

![](/files/-MNhtRQbPE4dpvx6hKo5)

Specify a descriptive title such as "Extensive Workflow Testing" and directory for the project to store its data.

3\. Create a new workspace for that project.

![](/files/-MNhtV17q3jbZ_xkFezy)

**Best practices:** Create a new workspace and run the Extensive Workflow in that workspace each time CryoSPARC updates and restarts. Name each workspace with the latest installed version of CryoSPARC that the job runs on. For example, when testing CryoSPARC v2.15.0, name the workspace "v2.15.0 Benchmark & Validation"

4\. Select the Job Builder from the sidebar and select the **"Extensive Workflow for T20S (BENCH) (BETA)"** job (under Workflows)

![](/files/-MNhtYqG36CI_p2G2AEt)

*5. (Optional)* If desired, change the workflow parameters.

Specifying a valid **"Movies data path"** and **"Gain reference path" is NOT required**; if the path does not exist on the system, cryoSPARC automatically downloads a \~8GB subset of the [T20S dataset](https://www.ebi.ac.uk/pdbe/emdb/empiar/entry/10025/) and deletes the download when the job finishes.

6\. Select "Queue" and choose a worker lane for the job, then select "Create"

![](/files/-MNhtjZ2u2uZg4ZeuU1s)

![](/files/-MNhtdCb2rzc8uWO36Uf)

After queuing, a modal opens with an overview of the workflow job progress. The job status should shortly change to "Running".

![](/files/-MNhtnOJYj5OGieN5nw-)

Close the modal with the `×` button. This shows a workspace overview with the child jobs that the Extensive Workflow job spawns to carry out T20S processing.

![](/files/-MNhtqv6zUBo4Zj8L_qy)

Once all child jobs successfully complete, the Extensive Workflow job status changes to "Completed". This means the installation was successful. Users may now be notified to start or resume processing!

## Troubleshooting Failed Jobs

If any child jobs fails, the extensive workflow times-out and its status is set to "Failed".

![](/files/-MNhu03VJpUkvMtbNyg3)

Scroll through the workspace to find other jobs with the "Failed" status. Open the job overview either by selecting the job number next to the status indicator (e.g., `J4`) , or by selecting the Job card and pressing the `Space` key.

Scroll to the bottom to see why the job failed.

![](/files/-MNhu3Bk0cZJRYwssn-U)

Example Import Movies failure because the Gain Reference was not found

Common failure reasons include

* [Cannot verify license](https://discuss.cryosparc.com/t/new-install-error-connecting-to-cryosparc-license-server/3825) or license key entered incorrectly
* [Incorrect filesystem permissions](https://discuss.cryosparc.com/t/file-r-w-permissions/3544)
* [CUDA not set up correctly](https://discuss.cryosparc.com/t/3d-variability-analysis-errors-v2-9-0/3110)
* [SSD cache full or not set up correctly](https://discuss.cryosparc.com/t/changing-ssd-no-run-directory/2916)
* [Worker registration issue](https://discuss.cryosparc.com/t/exit-status-255-error/2177)
* [Not enough GPU memory available](https://discuss.cryosparc.com/t/patchmotion-failure-2-13-2/3971)

Once the configuration issue is resolved, restart the Extensive Workflow job: Either create a new workspace and job as already noted, or clear the existing Extensive Workflow job and re-queue.

![](/files/-MNhuDBI4r-JWRUzvEf-)

![](/files/-MNhuGKprHt1qgpkjD1m)

## Additional Extensive Testing

For an even more extensive test of robustness, the Extensive Workflow job provides an advanced option called "Run all job types"

![](/files/-M7DHQLi2K0qV9AJqkFs)

![](/files/-M7DHQLj7ZUriqvSICWa)

1. Enable advanced mode near the top of the job builder
2. Enable to "Run all job types" switch

With this option enabled, the job runs additional child jobs in parallel. Use this to verify multi-GPU performance on a single node.

The following additional job types are included:

* Full-frame motion correction
* Global CTF estimation
* Local motion correction
* Multi-class *ab-initio* reconstruction
* Heterogeneous and non-uniform refinement
* 3D Variability

## Expected Results

Below are the results from our tests with CryoSPARC v3.1 on a 4GPU machine with the T20S subset.

On a machine with the [recommended system configuration](/setup-configuration-and-management/hardware-and-system-requirements#worker-node-cluster-worker-minimum-requirements), the Extensive Workflow takes \~1 hour with the default settings and \~1 hour 30 minutes with all job types enabled (note that some jobs run in parallel when enough GPUs are available).

| **Job Type**                             | Approximate Run Time (seconds) |
| ---------------------------------------- | ------------------------------ |
| **Import Movies**                        | 92                             |
| **Patch Motion Correction (Multi)**      | 220                            |
| **Full Frame Motion Correction (Multi)** | 75                             |
| **Patch CTF Estimation (Multi)**         | 66                             |
| **Curate Exposures**                     | 1.1                            |
| **Blob Picker**                          | 12                             |
| **Template Picker**                      | 13                             |
| **Inspect Picks**                        | 12                             |
| **Extract from Micrographs (CPU)**       | 39                             |
| **Extract from Micrographs (GPU)**       | 43                             |
| **Local Motion Correction**              | 180                            |
| **2D Classification**                    | 280                            |
| **Select 2D Classes**                    | 7.5                            |
| **Ab-Initio Reconstruction (1 class)**   | 450                            |
| **Ab-Initio Reconstruction (3 class)**   | 800                            |
| **Homogenous Refinement**                | 1940                           |
| **Heterogeneous Refinement (3 class)**   | 3000                           |
| **Non-Uniform Refinement**               | 4300                           |
| **Sharpen**                              | 32                             |
| **Validation**                           | 94                             |
| **Global CTF Refinement**                | 41                             |
| **Local CTF Refinement**                 | 46                             |
| **3D Variability**                       | 560                            |
| **3D Variability Display**               | 140                            |


# Guide: Performance Benchmarking (v4.3+)

This guide covers the new benchmarking tool in CryoSPARC that allows for benchmarking a worker’s filesystem, CPUs and GPUs. Available in CryoSPARC v4.3.0+.

## Overview

After installing CryoSPARC and verifying the instance is working correctly (see [Installation Testing](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/guide-installation-testing-with-cryosparcm-test)), use the Performance Benchmarking job to measure the performance of your system and compare it against references provided by Structura and your own past benchmarks.

The new “Benchmark” job is available in the CryoSPARC job builder and can be run on any worker lane connected to your CryoSPARC instance.

<figure><img src="/files/v32tYPEkJfmJtpwk2yMH" alt=""><figcaption></figcaption></figure>

The “Benchmark” job will make sure the benchmark data exists in the right location (and downloads it if it doesn’t), and runs the three benchmark tests in serial (CPU, Filesystem and GPU benchmarks) as specified.

{% hint style="warning" %}
CryoSPARC v5.0+ performance benchmark system is not backwards compatible with earlier versions. This means that performance benchmarks recorded in v5.0+ will be dropped when downgrading to v4.7 or below.

v5.0 instances include new updated reference performance benchmarks on bare metal and AWS node types.
{% endhint %}

### Benchmark Data

The benchmark data (17GB, compressed) is required to be downloaded and extracted into a location accessible by the job in order to run the benchmarks. As a convenience, this is automatically done by the Benchmark job when the required data does not exist in the project directory. The benchmark data can also be manually downloaded via the link provided below. Once manually downloaded and extracted, the absolute path to the folder can be specified in the “Benchmark Data Directory” parameter.

[Click here to download the benchmark data package directly from cloud storage.](https://s3.us-east-1.wasabisys.com/cryosparc-performance-benchmark-data/performance_benchmark_data_v1.tar.gz)

The benchmark data package contains movies, particles and volumes required for each of the tests. An abridged directory listing can be seen below:

```
.
├── class2D_test
│   └── maps.mrc
├── gpu_engine_test
│   ├── abinit_particles.cs
│   ├── abinit_volume.mrc
│   └── J586
│       └── extract
│           ├── 001411154804159785773_14sep05c_c_00003gr_00014sq_00005hl_00003es.frames_patch_aligned_doseweighted_particles.mrc
│           ├── (...)_14sep05c_c_00003gr_00014sq_00006hl_00003es.(...).mrc
│           ├── (...)_14sep05c_00024sq_00003hl_00005es.(...).mrc
│           ├── (...)_14sep05c_c_00003gr_00014sq_00007hl_00005es.(...).mrc
│           ├── (...)_14sep05c_c_00003gr_00014sq_00006hl_00005es.(...).mrc
│           ├── (...)_14sep05c_c_00003gr_00014sq_00005hl_00002es.(...).mrc
│           ├── (...)_14sep05c_c_00003gr_00014sq_00009hl_00004es.(...).mrc
│           ├── (...)_14sep05c_c_00003gr_00014sq_00011hl_00003es.(...).mrc
│           ├── (...)_14sep05c_00024sq_00004hl_00002es.(...).mrc
│           ├── (...)_14sep05c_c_00003gr_00014sq_00011hl_00002es.(...).mrc
│           ├── (...)_14sep05c_c_00003gr_00014sq_00008hl_00005es.(...).mrc
│           ├── (...)_14sep05c_c_00003gr_00014sq_00004hl_00004es.(...).mrc
│           ├── (...)_14sep05c_c_00003gr_00014sq_00010hl_00002es.(...).mrc
│           ├── (...)_14sep05c_c_00003gr_00014sq_00007hl_00004es.(...).mrc
│           ├── (...)_14sep05c_00024sq_00006hl_00003es.(...).mrc
│           ├── (...)_14sep05c_c_00003gr_00014sq_00005hl_00005es.(...).mrc
│           ├── (...)_14sep05c_00024sq_00003hl_00002es.(...).mrc
│           ├── (...)_14sep05c_c_00003gr_00014sq_00006hl_00002es.(...).mrc
│           ├── (...)_14sep05c_c_00003gr_00014sq_00002hl_00005es.(...).mrc
│           └── (...)_14sep05c_c_00003gr_00014sq_00011hl_00004es.(...).mrc
├── gpu_fsc_test
│   ├── half_map_A.mrc
│   └── half_map_B.mrc
├── movies
│   ├── eer
│   │   ├── FoilHole_2669035_Data_2668380_2668382_20200703_235726_Fractions.mrc.eer
│   │   ├── FoilHole_2669035_Data_2668383_2668385_20200703_235738_Fractions.mrc.eer
│   │   └── FoilHole_2669035_Data_2671097_2671099_20200703_235716_Fractions.mrc.eer
│   ├── mrc
│   │   ├── 17jul30a_b_00007gr_00002sq_v01_00002hl16_00002edhiii.frames.mrc
│   │   ├── 17jul30a_b_00007gr_00002sq_v01_00002hl16_00004edhiii.frames.mrc
│   │   └── 17jul30a_b_00014gr_00001sq_v01_00002hl16_00005edhiii.frames.mrc
│   └── tiff
│       ├── FoilHole_21044295_Data_21043958_21043960_20210422_050646_fractions.tiff
│       ├── FoilHole_21044296_Data_21043958_21043960_20210422_050719_fractions.tiff
│       └── FoilHole_21044297_Data_21043958_21043960_20210422_051105_fractions.tiff
└── picking_test
    ├── 014750136049583239150_14sep05c_00024sq_00004hl_00002es.frames_background.mrc
    ├── 014750136049583239150_14sep05c_00024sq_00004hl_00002es.frames_patch_aligned_ctf_spline.npy
    ├── 014750136049583239150_14sep05c_00024sq_00004hl_00002es.frames_patch_aligned_doseweighted.mrc
    ├── exposure_dataset_for_picking.cs
    ├── picking_templates.mrc
    └── templates_dataset_for_picking.cs
```

#### Particle Data Source

Particles previously processed in CryoSPARC from a subset of movies in EMPIAR-10025:

{% embed url="<https://www.ebi.ac.uk/empiar/EMPIAR-10025/>" %}

#### Movie Data Sources

**TIFF:** 3x K3 Super-resolution (11520, 8184) 70 frames: 1.16GB each from EMPIAR-10721:

{% embed url="<https://www.ebi.ac.uk/empiar/EMPIAR-10721/>" %}

**MRC:** 3x K2 (3710, 3838) 44 Frames 1.2GB each from EMPIAR-10249:

{% embed url="<https://www.ebi.ac.uk/empiar/EMPIAR-10249/>" %}

**EER:** 3x Falcon 4 (4096, 4096) 48 Frames: 500MB each from EMPIAR-10612:

{% embed url="<https://www.ebi.ac.uk/empiar/EMPIAR-10612/>" %}

### Sharing Data with Structura Biotechnology (Optional)

In each job, there is a parameter (”**Share benchmark data with Structura Biotechnology”,** disabled by default) to allow uploading of benchmark data to Structura’s servers. The data sent includes timings and hardware information, but does not include any user identifiable information.

An example of the data uploaded can be seen below:

```
{
    "type" : "gpu",
    "timings" : {
        "fsc_spherical" : {
            "put_mapr_on_gpu" : 0.5858473777771,
						...
        },
        "fsc_loose" : {
            "put_mapr_on_gpu" : 0.677032470703125,
						...
        },
				...
    },
    "cryosparc_version" : "v4.1.0",
    "created_at" : 1666985248.02303,
    "instance_information" : {
        "platform_node" : "server_hostname",
        "platform_release" : "4.15.0-142-generic",
        "platform_version" : "#146~16.04.1-Ubuntu SMP Tue Apr 13 09:27:15 UTC 2021",
        "platform_architecture" : "x86_64",
        "cpu_model" : "Intel(R) Xeon(R) CPU E5-1630 v4 @ 3.70GHz",
        "physical_cores" : 4,
        "total_memory" : "62.80GB",
        "available_memory" : "52.17GB",
        "used_memory" : "9.81GB",
        "ofd_soft_limit" : 1048576,
        "ofd_hard_limit" : 1048576
    },
    "job_params" : {
        "benchmark_data_dir" : null,
        "gpu_num_gpus" : 1,
        "send_data" : true,
        "test_random" : true,
        "test_sequential" : true,
        "use_all_gpus" : false,
        "use_ssd" : true
    },
    "gpu_name" : "NVIDIA GeForce GTX 1080 Ti",
    "gpu_bus_id" : "0000:02:00.0"
}
```

The data sent includes timings and hardware information, but does not include any user identifiable information. **Structura will use this data to maintain aggregate statistics about CryoSPARC performance in the wild and help us focus our optimization efforts on the jobs and codepaths with the most benefit to users. Users who do upload benchmark data should not expect any direct response from Structura.**

## Filesystem Benchmark

The filesystem benchmark employs a sequential read test for movies, and both a sequential and random read test for particles simulating real CryoSPARC workflows to benchmark the filesystem where the benchmark data exists.

Turn off the parameter “Use SSD for Tests” to disable the use of the caching system when performing the particle read tests. The job will instead report the time it takes to read the particles in a sequential and random pattern from the project directory instead of a local cache device.

### A Note About Filesystem Caching

On Linux, the “page cache” is an area of unused memory that is used to store data that the OS reads for later rapid retrieval. For example, when you read a 1GB file twice, the second access of the file will be faster, since the file blocks come directly from the cache in memory instead of the hard disk or SSD. The OS automatically frees up data stored in the page cache as more memory is requested by other applications. The Benchmark job attempts to drop files that it uses from the page cache by using the [`posix_fadvise`](https://linux.die.net/man/2/posix_fadvise) function to declare that the files “will not be accessed in the near future” (`POSIX_FADV_DONTNEED`). Doing so allows subsequent runs of the Benchmark job to be reproducible (meaning that the numbers reported by the job won’t be skewed by faster read times) without having to manually drop the page cache.

{% hint style="info" %}
**Dropping the page cache:** The following two commands first instruct the kernel to [write dirty pages to disk](https://linux.die.net/man/8/sync), then [drop the page cache](https://www.kernel.org/doc/Documentation/sysctl/vm.txt):

`sudo bash -c 'sync; echo 1 > /proc/sys/vm/drop_caches'`
{% endhint %}

Note that there still may be other caches in play if your data is hosted on other machines (e.g., a storage cluster’s cache).

### Sequential Read Test - Movies

To benchmark sequential reading, which is relevant in the early stages of data processing, three different types of movies (TIFF, EER and MRC) are read and timed. The test reports the averages of the total I/O time taken. This measures the performance of the storage volume on which the movies are located. To benchmark a storage volume that is different from the project directory, copy the benchmark data to the new location and specify it in the “**Benchmark Data Directory”** parameter. For more information on the sources of each of the movies, see [Benchmark Data](#benchmark-data).

For TIFF and EER movies, only the time it takes for the system to read the data into memory is recorded. The time it takes for the movies to be decompressed (which is always performed when reading TIFF and EER movies) is timed but not recorded.

### Sequential Read Test - Particles

To benchmark sequential particle reads, which are relevant during some parts of particle processing, a small particle stack (50,000 particles with shape (256,256) across 500 files) is randomly generated using [`numpy.random.randn`](https://numpy.org/doc/stable/reference/random/generated/numpy.random.randn.html), written to the project directory, cached onto the cache device (if available and enabled), then read (in a sequential pattern) back into memory. The time it takes the system to cache the particles (`particle_cache_time`), read the particles sequentially (`particle_sequential_read_time`) and the rate at which the particles are read (`particle_sequential_read_rate`) are recorded.

### Random Read Test - Particles

To benchmark random reads, which are relevant during most parts of particle processing, the same particle stack created during the sequential read test is used, but this time the particles are read in a random pattern. The time it takes the system to read the particles randomly (`particle_random_read_time`) and the rate at which the particles are read (`particle_random_read_rate`) are recorded.

## CPU Benchmark

The CPU Benchmark reads the same TIFF and EER movies from the Filesystem test (see [Sequential Read Test](#sequential-read-test-movies)), but instead reports the time it takes to decompress the movies, which is heavily dependent on CPU and Memory performance.

<figure><img src="/files/K8STzarbDTQM7emp9swy" alt=""><figcaption></figcaption></figure>

Note that in order to measure decompression time, the CryoSPARC environment variable `CRYOSPARC_TIFF_IO_SHM` must be set and turned on (which by default, it is). [See Environment Variables](https://guide.cryosparc.com/setup-configuration-and-management/management-and-monitoring/environment-variables#cryosparc_worker-config.sh). This parameter tells the IO system to first copy the contents of TIFF and EER files to `/dev/shm` (a temporary file storage system backed by RAM) before decompressing it, allowing the system to distinguish IO time from decompression time. Note that this parameter also increases performance on some networked file systems.

Decompression time is averaged from three runs of different movies and reported as `tiff` and `eer` in the CPU tab of the Benchmark viewer.

## GPU Benchmark

The GPU benchmark executes a collection of functions from CryoSPARC jobs on each of the worker’s GPUs (unless the “**Number of GPUs to benchmark**” parameter is specified, in which case only the specified number of GPUs are benchmarked), and times them. The tests include:

1. FSC Calculations using different masks:
   * Spherical Mask (`fsc_spherical`)
   * Loose Mask (`fsc_loose`)
   * Tight Mask (`fsc_tight`)
   * Noise Sub Mask (`fsc_noisesub`)
2. Non-Uniform Refinement’s core algorithm (`matched_cv_filter_estimation`)
3. Particle picking’s core algorithm (`picking`)
4. CryoSPARC’s core alignment and reconstruction algorithm (”Engine”), tested with various parameter combinations:
   * Particles in cache
     * Using 1 CPU thread
       * Using a trilinear interpolation kernel
         * Using C1 symmetry
           * Using Pose Maximization
             * Test A (`disk_single_linear10_max_C1`)
       * Using a tricubic interpolation kernel
         * Using C1 symmetry
           * Using Pose Maximization
             * Test B (`disk_single_linear20_max_C1`)
     * Using 2 CPU threads
       * Using a trilinear interpolation kernel
         * Using C1 symmetry
           * Using Pose Maximization
             * Test C (`disk_multi_linear10_max_C1`)
       * Using a tricubic interpolation kernel
         * Using C1 symmetry
           * Using Pose Maximization
             * Test D (`disk_multi_linear20_max_C1`)
   * Particles in memory
     * Using 1 CPU thread
       * Using a trilinear interpolation kernel
         * Using C1 symmetry
           * Using Pose Maximization
             * Test E (`memory_single_linear10_max_C1`)
           * Using Pose Marginalization
             * Test F (`memory_single_linear10_marg_C1`)
         * Using D7 symmetry
           * Using Pose Maximization
             * Test G (`memory_single_linear10_max_D7`)
       * Using a tricubic interpolation kernel
         * Using C1 symmetry
           * Using Pose Maximization
             * Test H (`memory_single_linear20_max_C1`)
           * Using Pose Marginalization
             * Test I (`memory_single_linear20_marg_C1`)

{% hint style="info" %}
As of CryoSPARC v4.4+, `memory_multi_*` (particles in memory + multithreaded) tests have been removed.
{% endhint %}

#### Non-Uniform Refinement’s core algorithm

The core algorithm used in Non-Uniform Refinement performs multiple data transfers to and from the GPU, while performing hundreds of GPU-accelerated FFTs. This test stresses the memory performance of the GPU and PCIe bandwidth of the CPU, and is limited by the performance of a single CPU core.

#### CryoSPARC’s core reconstruction algorithm

The different parameter combinations specified for the core reconstruction algorithm tests code paths used by various CryoSPARC jobs including Homogeneous Refinement, Non Uniform Refinement, 3D Classification, and more. The test name (e.g., `memory_multi_linear20_marg_C1`) is comprised of the following parameters used to perform the test:

\<particle location>*\<number of CPU threads>*\<interpolation kernel>*\<pose assignment method>*\<symmetry operator>

* Particle Location:
  * Particles can either be stored on the cache device (SSD) if caching is enabled, or read into memory. When particles are in memory, IO time becomes negligible.
* Number of CPU Threads:
  * CryoSPARC’s core algorithm can be run with one or two threads. Most of the time, it’s run with two threads, so that particle IO and GPU computation is performed concurrently. Note that in tests using 2 CPU threads, some timing numbers are not accurate due to the concurrency of the computation, which is why only the `overall` time is reported. In these cases, it’s best to compare the timings from the corresponding single-threaded test.
* Interpolation Kernel:
  * CryoSPARC’s refinement and classification algorithms use two main interpolation kernels: trilinear (`linear10`) and tricubic (`linear20`) to interpolate values of the 3D density in Fourier space. Interpolation is necessary when rotating and projecting the 3D density, which is used in the orientation search step in most refinement/classification/variability jobs.

    Trilinear interpolation is significantly less computationally expensive than tricubic interpolation, requiring only 8 array accesses (vs. 64) of the underlying 3D density. Trilinear interpolation is also hardware-accelerated on NVIDIA GPUs through CUDA, whereas tricubic interpolation is not.
* Pose Assignment Method:
  * Non-Uniform Refinement supports either pose “`max`imization” or “`marg`inalization” during the reconstruction of the 3D density from the particle images. Maximization means that each particle is assigned a single 3D pose and shift during reconstruction. Alternatively, marginalization allows each particle to be assigned multiple 3D poses and shifts, each being weighted by their relative likelihoods under the image formation model. Maximization is usually sufficient, but for small particles or noisy datasets, marginalization helps to account for uncertainty in estimating the poses. When reconstructing the 3D density, maximization only has to insert each particle image into the 3D reconstruction **once;** marginalization is more computationally expensive because it requires inserting each image into the reconstruction multiple times.
* Symmetry Operator:
  * CryoSPARC’s core reconstruction algorithm supports many symmetry operators, but C1 and D7 were chosen for these benchmarks as a way to turn “off” and “on” the code paths respectively that enable symmetry.

## Interpreting Results Using The Benchmark Viewer

To view and compare previous Benchmark results and reference benchmarks provided by Structura, navigate to the “Benchmarks” tab inside the “Manage” panel.

<figure><img src="/files/suIvaRUOtO9vnu8eDYPJ" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/IQhOD2K811MWF7bgotNS" alt=""><figcaption></figcaption></figure>

Under each sub-tab (CPU, File System, GPU, Extensive Validation), there will be reference benchmarks provided by Structura which can be used as a comparison against benchmarks run on the current instance.

To compare multiple references, select them from the table using the checkboxes and click the “Compare” button on the top right side of the screen.

<figure><img src="/files/pSIKXA4BgbPaIplPx9nU" alt=""><figcaption></figcaption></figure>

In the comparison view, each benchmark is a column, and their timings are listed as rows. The overall time that the benchmark took is listed under the “Time” sub-column (**A**), and the portion of how long it took relative to the other timings is represented as a percentage in the “Pct” sub-column.

When a benchmark (column) is selected, it becomes the base “Reference” *B1* for the “Speedup” columns *B2*, which are available for all other benchmarks in the comparison view. The “Speedup” is calculated as

$$
Reference(seconds)/Current(seconds),
$$

which helps to easily glean how much faster or slower a timing is *in comparison to* the reference.

When a timing is hovered over, its details will be displayed in the “Benchmark Details” section on the right side of the page (**C**).

<figure><img src="/files/6pyW4iwNuURBrrf3Wp1R" alt=""><figcaption></figcaption></figure>

To view more detailed timings (available for the GPU benchmark only), click on the “+” button to expand a row (**D**). These sub-timings are the low-level functions that get called inside of CryoSPARC’s core reconstruction algorithm. The “Tags” column (**E**) indicates what hardware component each function’s speed is most dependent on.

For example, for the `setup_scales` sub-timing, the relevant component tags are “PCIe Latency/Bandwidth Speed” and “GPU/CPU Memory Allocation” because the function allocates space for a float32 array in CPU and GPU memory, fills it with data in CPU Memory, then downloads the contents of the array from CPU memory to it’s corresponding location on GPU memory. When a “download” happens, this occurs over the PCIe lanes that connect the CPU to the GPU, where the link speed (determined by e.g., PCIe Gen. 3 on most GPUs and PCIe Gen. 4 on NVIDIA Ampere and Ada architectures) matters the most in determining how fast this happens.

### Component Tags

<table><thead><tr><th width="297">Tag Name</th><th>Most likely bottleneck</th></tr></thead><tbody><tr><td>CPU Performance</td><td>- single core clock speed</td></tr><tr><td>GPU Performance</td><td>- float32 performance and memory bandwidth</td></tr><tr><td>PCIe Latency/Bandwidth Speed</td><td>- PCIe generation (3, 4) and number of lanes per slot (x8, x16)</td></tr><tr><td>GPU/CPU Memory Allocation</td><td>- general cpu/gpu performance</td></tr><tr><td>Input/Output Speed</td><td>- random read speeds of the storage device where particle images are located</td></tr></tbody></table>

### Raw Data

At the end of the benchmark job, results are saved as a JSON and CSV in the job directory. The exact path of the files can be seen at the end of each test in the job’s Event Log.

```
Writing benchmark data to /bulk9/data/dev_projects/CS-peformance-benchmark/J73/J73_fs_benchmark_data.json

Writing benchmark data to /bulk9/data/dev_projects/CS-peformance-benchmark/J73/J73_fs_benchmark_data.csv
```

To view the original Benchmark job that a benchmark was created from, right click on the column header and select “Show job in sidebar”. The JSON and CSV results can also be downloaded from this context menu.

<figure><img src="/files/GKXz2KBfnpU9d4d8Rbgk" alt=""><figcaption></figcaption></figure>

## Performance Benchmarking Entire Jobs with the Extensive Validation Job

The Extensive Workflow job is now called the Extensive Validation job (v4.3.0+).

The Extensive Validation job is a job that creates and queues other jobs in a pre-defined workflow. Workflows available are for the [EMPIAR-10025](https://www.ebi.ac.uk/empiar/EMPIAR-10025/) and [EMPIAR-10305](https://www.ebi.ac.uk/empiar/EMPIAR-10305/) datasets, which are downloaded when the job is run. If you run the Extensive Validation job in "Benchmark" mode, each job defined in the workflow will run in sequence. This will allow you to compare the overall performance of each job in the Benchmark UI, along with the CPU, Filesystem, and GPU performance benchmarks.

First, create an “Extensive Validation” job and select “Benchmark” as the value for the “Run Mode” parameter:

<figure><img src="/files/Eu7lgVvdwpNbIhWYZecp" alt=""><figcaption></figcaption></figure>

In “Benchmark Mode”, jobs that support multi-GPU parallelization (such as Patch Motion Correction, Patch CTF Estimation, and 2D Classification) can be allocated multiple GPUs. To allocate multiple GPUs, specify a number greater than 1 for the “Number of GPUs to use” parameter field, and either select a lane or specify the exact GPUs using the “Run on specific GPUs” tab in the Resource Selection panel.

<figure><img src="/files/fXvQFlTNgJw8GPaXnIo5" alt=""><figcaption></figcaption></figure>

For more information on the jobs that are launched by the Extensive Validation job in benchmark mode, see the Extensive Validation documentation here:

{% content-ref url="/pages/-MNhrOW1ytZ8-U3WkV0r" %}
[Guide: Verify CryoSPARC Installation with the Extensive Validation Job (v4.3+)](/setup-configuration-and-management/software-system-guides/tutorial-verify-cryosparc-installation-with-the-extensive-workflow-sysadmin-guide)
{% endcontent-ref %}

## Appendix

### Drop the page cache

First, instruct the kernel to [write dirty pages to disk](https://linux.die.net/man/8/sync), then [drop the page cache](https://www.kernel.org/doc/Documentation/sysctl/vm.txt):

`sudo bash -c 'sync; echo 1 > /proc/sys/vm/drop_caches'`


# Guide: Download Error Reports

How to download job and system-level error reports from within the application.

{% hint style="warning" %}
The information in this section applies to CryoSPARC v4.0+.
{% endhint %}

## Error Reporting

### Job Error Report

A complete job error report including job details, job event logs, browser diagnostics, and system logs can be downloaded from the Event Log tab of the job inspection dialog.

When downloading the job error report, there is an option to include event log images. You may wish to omit the images in case they are too large and take too long to download or transfer, or if you do not wish to include images of your data/results when sharing error reports.

System logs up to one week prior will be included in the download.

The job error report can be used to further diagnose a job failure when the event log and job log aren’t showing the cause, or as a convenient way to share error information with a system administrator.

<figure><img src="/files/XcKeL7vMbO0o2on1sqZX" alt=""><figcaption></figcaption></figure>

Example job report contents:

<figure><img src="/files/JkCCOt47qu5E9wIphXXc" alt=""><figcaption></figcaption></figure>

### System Error Report

A system error report including instance information, browser diagnostics, and system logs can be downloaded from the Instance Logs tab of the Admin panel. System logs up to one week prior will be included in the download.

The system error report is useful for administrators to diagnose the cause of system-wide errors, and as a convenient way to share error information on the [CryoSPARC Discussion Forum](https://discuss.cryosparc.com/) or CryoSPARC support.

<figure><img src="/files/sIqgHczXMxCpCk4DlCOJ" alt=""><figcaption></figcaption></figure>

Example system error report contents:

<figure><img src="/files/T81NiS8OsDz3BM1KezDM" alt=""><figcaption></figcaption></figure>


# Guide: Maintenance Mode and Configurable User Facing Messages

Pause the job queue during updates and set optional user facing messages.

{% hint style="warning" %}
The information in this section applies to CryoSPARC v4.0+.
{% endhint %}

## Maintenance Mode

CryoSPARC’s maintenance mode can be used to pause the job queue so that running jobs complete, but queued jobs are not launched. This allows a busy CryoSPARC instance to be gracefully shut down, re-started, or updated without causing jobs to be terminated or to fail. For example, maintenance mode can be turned on when updating a CryoSPARC instance to prevent new jobs from starting until after the update is rolled out. This can help provide a smooth transition for users during an update by allowing running jobs to complete but prevent new jobs from running until the update is successfully installed and maintenance mode is turned off.

Maintenance mode stops worker nodes from starting new jobs while allowing them to finishing any jobs already in progress. Users can continue to use the application to browse results and to create and queue jobs, but those jobs will only begin processing after maintenance mode is turned off. Users are shown a message in the user interface indicating that maintenance mode is on and what it means for them.

Once an administrator has turned on maintenance mode, the queue will be paused. When existing running jobs are complete, the instance will be idle and the administrator can re-start, update, etc. Once this is complete, the administrator can turn off maintenance mode and previously queued jobs will be launched and processing will continue.

<figure><img src="/files/P2vdrRe60gJ7Nr3nyVIx" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
Turning on maintenance mode does not automatically pause running CryoSPARC Live sessions. Please pause any running Live sessions before updating CryoSPARC.
{% endhint %}

Maintenance mode can be managed using `cryosparcm`.

To enable maintenance mode:

```bash
$ cryosparcm maintenancemode on

Turning maintenance mode ON
```

To check the status of maintenance mode:

```bash
$ cryosparcm maintenancemode status

Maintenance mode is currently ON
```

To turn maintenance mode off:

```bash
$ cryosparcm maintenancemode off

Turning maintenance mode OFF
```

## Message of the Day

In CryoSPARC v4.0+, administrators can set a message that displays as a banner at the top of the dashboard in the UI. This may be useful for announcing scheduled downtimes, planned maintenance, or other information that users should be aware of. The message banner has a customizable title and body, and can be toggled on and off.

<figure><img src="/files/pu7a4PeTboQqD4EseIsW" alt=""><figcaption></figcaption></figure>

he message banner can be managed using `cryosparcm`.

To set the banner title and message:

```bash
$ cryosparcm cli "set_instance_banner(True, 'This is the message title', 'This is the message body')"

{'active': True, 'body': 'This is the message body', 'title': 'This is the message of the day'}
```

To get the current banner status, title, and message:

```bash
$ cryosparcm cli "get_instance_banner()"

{'active': True, 'body': 'This is the message body', 'title': 'This is the message of the day'}
```

To toggle the banner on:

```bash
$ cryosparcm cli "set_instance_banner(True)"

{'active': True, 'body': 'This is the message body', 'title': 'This is the message of the day'}
```

To toggle the banner off:

```bash
$ cryosparcm cli "set_instance_banner(False)"

{'active': False, 'body': 'This is the message body', 'title': 'This is the message of the day'}
```

## Login Message

In v4.0+, CryoSPARC administrators can set a message that displays as modal the first time a user logs into CryoSPARC. This may be useful for announcing important information that users should be aware of when logging in for the first time. The message banner has a customizable title and body, and can be toggled on and off.

<figure><img src="/files/bZfmnP5V7ot4oP48RZWM" alt=""><figcaption></figcaption></figure>

The message banner can be managed using `cryosparcm`.

To set the banner title and message:

```bash
$ cryosparcm cli "set_login_message(True, 'This is the login message', 'This is the message body')"

{'active': True, 'body': 'This is the message body', 'title': 'This is the login message'}
```

To get the current banner status, title, and message:

```bash
$ cryosparcm cli "get_login_message()"

{'active': True, 'body': 'This is the message body', 'title': 'This is the login message'}
```

To toggle the banner on:

```bash
$ cryosparcm cli "set_login_message(True)"

{'active': True, 'body': 'This is the message body', 'title': 'This is the login message'}
```

To toggle the banner off:

```bash
$ cryosparcm cli "set_login_message(False)"

{'active': True, 'body': 'This is the message body', 'title': 'This is the login message'}
```


# Guide: User Management

User creation, management, setting roles and password management through the CryoSPARC user interface.

{% hint style="warning" %}
The information in this section applies to CryoSPARC ≤v3.3. For CryoSPARC v4.0+, please see: [Admin Panel](/application-guide/admin-panel)
{% endhint %}

CryoSPARC v2.12+ offers UI-based user management tools to quickly create and change the roles of users in your instance. Additionally, it's possible for a user to reset their password with the help of an admin user, directly through the UI. To learn more about how to accomplish these tasks, read on.

* **Create users** through the admin page in cryoSPARC interface
* **Manage existing users** by promoting them to admin or converting admins to regular users
* **The reset password** functionality allows users to request and set a new password without needing an administrator to use the command-line interface

## Create a New User

1.Enter the user management page by clicking on your user name and selecting 'Admin' in the menu. You must be an admin user in order to access the user management section. When installing cryoSPARC, the first user that is created through the command-line interface is set as an admin.

![](/files/-MNeet3qB8Yg-b8BGDMU)

2\. Fill the 'Add a New User' form at the bottom of the table with the new user's email address and username, along with their first and last name.

![](/files/-MNeedAmnfRX3HcfwIzd)

3\. The new user will be added to the table of users. In the 'Tokens' column, you can click the created user's 4-digit registration token and copy it to the clipboard, sending it to them (email, instant message, etc.)

![](/files/-MNeeQlQVgZSxcbPrdUs)

4\. Now that your new user has a registration token, they can click the 'New Account' link on the cryoSPARC login form and enter their email address, registration token and new password.

![](/files/-MNeemYkcBRtgJM5ySiY)

![](/files/-MNeeX2eXo-b0W8ICFwU)

5\. The new user will automatically be logged-in when they complete the account creation process.

## Change a User's Role

1.Enter the user management page by clicking on your user name and selecting 'Admin' in the menu.

![](/files/-MNeet3qB8Yg-b8BGDMU)

![](/files/-MNeedAmnfRX3HcfwIzd)

2\. In the 'Role' column, click the role of a particular user and confirm you wish to change their status. Users who are not admin will be promoted to admin status and existing admin users will be converted into normal users.

![](/files/-MNeehW04Y64gpM_CyPK)

## Reset your Password

{% hint style="info" %}
Structura Biotechnology cannot reset your password. Only your local CryoSPARC administrator can reset your password.
{% endhint %}

1. The user wishing to reset their password should select the 'Reset Password' link on the cryoSPARC login form and enter their email address ('I need a reset token' option selected).

![](/files/-MNeemYkcBRtgJM5ySiY)

![](/files/-MNeepLWbyYQ-J6s1MVH)

2\. An existing admin user can access the reset token via the user management page (click on your user name and select 'Admin' in the menu).

![](/files/-MNeet3qB8Yg-b8BGDMU)

3\. The admin user should retrieve the 4-digit reset token and securely communicate it to the user in need of a password reset (clicking the token will copy it to the clipboard).

![](/files/-MNeevvZuGLFhxPlQQUS)

4\. The user in need of a password reset fills out the *I have a reset token* portion of the form with e-mail address, the reset token from the previous step and the new password and will automatically be logged-in after clicking the *Reset Password* button.

![](/files/-MNef5kutPGH-ydyRUxz)


# Guide: Multi-user Unix Permissions and Data Access Control

Tips on how to manage permissions and data access control.

{% hint style="warning" %}
This guide explains how to set up accounts when there are multiple unix users interacting with the same CryoSPARC instance. For standard instructions about how to set up unix accounts for CryoSPARC installation see the [Installation Pre-requisites](https://guide.cryosparc.com/setup-configuration-and-management/cryosparc-installation-prerequisites#2.-common-unix-user-account).
{% endhint %}

The permissions system built into Linux (and all Unix-like operating systems) can be used to make it easier for multiple users to interact with CryoSPARC's data, or to establish a degree of separation between data belonging to different research teams. This page offers some tips on how this can be done.

{% hint style="info" %}
Changing file permissions has security implications. If you aren't familiar with how Unix permissions work, it is advisable to consult with your system administrator, consult your operating system's documentation, or do some online research into Unix permissions before running the commands in this guide.
{% endhint %}

## Allowing users' Linux accounts to access CryoSPARC files

CryoSPARC should be installed under its own operating system account. Typically, individual researchers will have their own accounts for accessing the computer(s) that CryoSPARC is installed on. By default on most Linux distributions, user accounts cannot modify files created by other user accounts, so the accounts belonging to individual researchers can only read files created or owned by CryoSPARC. This can be a barrier when trying to, for example, copy exported job files into another project for import. The standard permission system built into Linux can be used to work around this problem. Below is an example of how.

For the purpose of this example, we'll assume we are working with a fresh CryoSPARC installation. We will create a cryosparc user account (all CryoSPARC processes will run as this user), and we'll create two user accounts for people who will be using CryoSPARC.

```
root@host:~# useradd cryosparc
root@host:~# useradd alice
root@host:~# useradd tom
```

We'll add Alice and Tom's accounts to the "cryosparc" group:

```bash
root@host:~# usermod -aG cryosparc alice
root@host:~# usermod -aG cryosparc tom
```

Since in this example we're assuming this is a brand new CryoSPARC installation, we'll set up the directory that we will be storing our project files in. We'll then change the ownership of that directory so that it's under the control of the cryosparc account and group, and lastly we'll change the permissions on that directory. Specifically, we want `g+ws`, which will make anyone in the cryosparc group able to write to that directory (that's the `w`), and will make it so that any time new file is created in that directory, the file is owned by the cryosparc group (that's the `s`).

```bash
root@host:~# mkdir -p /data/cryosp_projs
root@host:~# chown cryosparc:cryosparc /data/cryosp_projs
root@host:~# chmod g+ws /data/cryosp_projs/
```

A demonstration of the result:

```
# the cryosparc user creates a bunch of files

root@host:~# su cryosparc
cryosparc@host:/root$ touch /data/cryosp_projs/file1
cryosparc@host:/root$ touch /data/cryosp_projs/file2
cryosparc@host:/root$ touch /data/cryosp_projs/file3
cryosparc@host:/root$ ls -l /data/cryosp_projs/
total 0
-rw-rw-r-- 1 cryosparc cryosparc 0 Jun 18 17:05 file1
-rw-rw-r-- 1 cryosparc cryosparc 0 Jun 18 17:05 file2
-rw-rw-r-- 1 cryosparc cryosparc 0 Jun 18 17:05 file3
cryosparc@host:/root$ exit


# alice logs in, and also creates a file in that directory

root@host:~# su alice
alice@host:/root$ touch /data/cryosp_projs/file4
alice@host:/root$ ls -l /data/cryosp_projs/
total 0
-rw-rw-r-- 1 cryosparc cryosparc 0 Jun 18 17:05 file1
-rw-rw-r-- 1 cryosparc cryosparc 0 Jun 18 17:05 file2
-rw-rw-r-- 1 cryosparc cryosparc 0 Jun 18 17:05 file3
-rw-rw-r-- 1 alice     cryosparc 0 Jun 18 17:06 file4
alice@host:/root$ exit

# notice how the file alice made is owned by the cryosparc group.
#
# now tom can log in and is able to modify the files created by...
# ... both alice and cyosparc

root@host:~# su tom
tom@host:/root$ rm /data/cryosp_projs/file2 
tom@host:/root$ rm /data/cryosp_projs/file4
tom@host:/root$ ls -l /data/cryosp_projs/
total 0
-rw-rw-r-- 1 cryosparc cryosparc 0 Jun 18 17:05 file1
-rw-rw-r-- 1 cryosparc cryosparc 0 Jun 18 17:05 file3
tom@host:/root$ touch /data/cryosp_projs/file
tom@host:/root$ ls /data/cryosp_projs/ -l
total 0
-rw-rw-r-- 1 tom       cryosparc 0 Jun 18 17:07 file
-rw-rw-r-- 1 cryosparc cryosparc 0 Jun 18 17:05 file1
-rw-rw-r-- 1 cryosparc cryosparc 0 Jun 18 17:05 file3
```

{% hint style="info" %}
On some systems, the default `umask` variable may disable group write permissions. This would cause new files created by a user to not have group write permission unless manually changed with `chmod`. You can determine what the current umask is by running `umask` at a command prompt. If the second-to-last digit is not zero, you may run into issues. Putting `umask 0002` in the \~/.bashrc file for the cryosparc account and each individual user account will correct this issue if it occurs.
{% endhint %}

## Establishing teams of users, and limiting access to data owned by another team

Another area where unix permissions can help is in separating CryoSPARC users into teams that cannot interact with each other's data. Be sure to read the above section before this one, as a few important ideas are not repeated here.

{% hint style="info" %}
While this guide can be used as a basis for limiting access to data and projects at the level of unix user accounts, the cryosparc user can still access the data from all teams. This is necessary for CryoSPARC to function correctly. As a result, users could still access another group's work by using the CryoSPARC user interface and navigating to a project directory owned by another group. One group could, for example, import raw data owned by another group. This will be addressed in a future CryoSPARC release.
{% endhint %}

This guide will proceed by example, similar to the previous.

We assume we're starting with a blank slate. Suppose we have four researchers in two separate research groups: Alice and Tom are in the 'lab1' group, Dmitri and Sonja are in the 'lab2' group.

```
root@host:~# useradd tom
root@host:~# useradd alice
root@host:~# useradd dmitri
root@host:~# useradd sonja
root@host:~# groupadd lab1
root@host:~# groupadd lab2
root@host:~# usermod -aG lab1 tom
root@host:~# usermod -aG lab1 alice
root@host:~# usermod -aG lab2 dmitri
root@host:~# usermod -aG lab2 sonja 

# add the cryosparc user account to both lab1 and lab2

root@host:~# useradd cryosparc
root@host:~# usermod -aG lab1 cryosparc
root@host:~# usermod -aG lab2 cryosparc
```

Create the directory that will be used to store CryoSPARC projects, and establish the same permissions as in the previous guide.

To briefly recap, we're going to use "chmod g+sw", where the "s" means that files created within the folder will be associated with the folder's group. Right now, that isn't terribly useful as we haven't added any users to the cryosparc group. But that "s" setting - the "setgid" bit, will be set on any created subdirectories as well, which will become important shortly.

```
root@host:~# mkdir -p /data/cryosp_projs
root@host:~# chmod g+ws /data/cryosp_projs/
root@host:~# chown cryosparc:cryosparc /data/cryosp_projs/
```

At this point, create some projects via the CryoSPARC UI which, will automatically create project subfolders. (e.g. "P1", "P2", "P3", etc). Decide which project subfolders should be owned by which group, and change the folder ownership appropriately as shown below. Note that the 'mkdir' lines are just simulating what the UI will do when creating a project - these don't need to actually be run.

```
root@host:~# su cryosparc
cryosparc@host:/root$ cd /data/cryosp_projs/
cryosparc@host:/data/cryosp_projs$ mkdir P1
cryosparc@host:/data/cryosp_projs$ mkdir P2
cryosparc@host:/data/cryosp_projs$ mkdir P3
cryosparc@host:/data/cryosp_projs$ chgrp lab1 P1
cryosparc@host:/data/cryosp_projs$ chgrp lab2 P2
cryosparc@host:/data/cryosp_projs$ chgrp lab2 P3
cryosparc@host:/data/cryosp_projs$ chmod o-rx *
cryosparc@host:/data/cryosp_projs$ ls -l
total 12
drwxrws--- 2 cryosparc lab1 4096 Jun 18 17:12 P1
drwxrws--- 2 cryosparc lab2 4096 Jun 18 17:12 P2
drwxrws--- 2 cryosparc lab2 4096 Jun 18 17:12 P3
cryosparc@host:/data/cryosp_projs$ exit
```

Take a look at the output of the `ls -l` command. Notice that each project folder has an "s" where the group "x" would normally be. This happened because we added the "s" bit to the `cryosparc_projs` directory before we created the projects (it could be added with chmod if we were doing this retroactively). As a result of the presence of that "s" bit and the fact that each project folder is now owned by one of the lab groups, any files created inside the project folders will be owned by the corresponding lab group, which allows our researcher accounts to access and modify them as appropriate.

Notice also that we used `chmod o-rx`, meaning that users who aren't cryosparc or part of the appropriate lab group cannot even see the contents of the project directories.

The resulting configuration is demonstrated below:

```
root@host:~# su tom
tom@host:/root$ cd /data/cryosp_projs/
tom@host:/data/cryosp_projs$ ls
P1  P2	P3
tom@host:/data/cryosp_projs$ cd P1/
tom@host:/data/cryosp_projs/P1$ touch file1
tom@host:/data/cryosp_projs/P1$ cd ..
tom@host:/data/cryosp_projs$ cd P2
bash: cd: P2: Permission denied
tom@host:/data/cryosp_projs$ exit


root@host:~# su sonja
sonja@host:/root$ cd /data/cryosp_projs/
sonja@host:/data/cryosp_projs$ ls P1
ls: cannot open directory 'P1': Permission denied
sonja@host:/data/cryosp_projs$ touch P3/file
sonja@host:/data/cryosp_projs$ ls -l P3
total 0
-rw-rw-r-- 1 sonja lab2 0 Jun 18 17:14 file
sonja@host:/data/cryosp_projs$ exit


root@host:~# su cryosparc
cryosparc@host:/root$ cd /data/cryosp_projs/
cryosparc@host:/data/cryosp_projs$ ls
P1  P2	P3
cryosparc@host:/data/cryosp_projs$ find
.
./P3
./P3/file
./P1
./P1/file1
./P2
cryosparc@host:/data/cryosp_projs$ rm P3/file
cryosparc@host:/data/cryosp_projs$ rm P1/file1
```


# Guide: Lane Assignments and Restrictions

Assigning CryoSPARC users to specific scheduler lanes.

{% hint style="warning" %}
The information in this section applies to CryoSPARC v4.1+.
{% endhint %}

## Lane Assignments and Restrictions

CryoSPARC users can be assigned access to specific lanes in the CryoSPARC scheduler. Having access to a lane means that the user can queue jobs to the lane and use its resources. Users not assigned to a lane do not have access to it and cannot queue jobs on that lane. By default, when a CryoSPARC user is created, they will be assigned all existing lanes. Similarly, when a new lane is added, all users will be assigned to that lane.

A user’s lane assignments can be viewed and modified using the “Lane Restrictions” tab of the admin panel in the UI.

<figure><img src="/files/DEZW1nzJdz8HXmx6dHKs" alt=""><figcaption></figcaption></figure>

Lanes can be assigned/unassigned for each user by checking the boxes next to the lane names and clicking the transfer arrow in the corresponding direction.

To get a user's lane assignments using the CLI:

{% tabs %}
{% tab title="v5" %}

```python
$ cryosparcm cli "api.users.get_lanes('user123@structura.bio')"
['cryoem1', 'cryoem2', 'cryoem3']
```

{% endtab %}

{% tab title="v4" %}

```javascript
$ cryosparcm cli "get_user_lanes('user123@structura.bio')"
['cryoem1', 'cryoem2', 'cryoem3']
```

{% endtab %}
{% endtabs %}

To modify a user’s lane assignments using the CLI:

{% tabs %}
{% tab title="V5" %}

```javascript
$ cryosparcm cli "api.users.set_lanes('user123@structura.bio', ['cryoem1'])"
```

{% endtab %}

{% tab title="V4" %}

```python
$ cryosparcm cli "set_user_lanes('user123@structura.bio', ['cryoem1'])"
```

{% endtab %}
{% endtabs %}


# Guide: Priority Job Queuing

How to prioritize jobs to override the CryoSPARC scheduler's default behaviour.

## Overview

As of v3.0.0+, CryoSPARC jobs have a `priority` property that the CryoSPARC scheduler uses to sort jobs to be executed. **The `priority` is sorted in descending order, meaning jobs with higher priority will get queued first.** When the scheduler encounters jobs that have the same priority, it will sort the jobs based on the time they entered the queue in ascending order.

![The priority of a job can be seen in the resource manager.](/files/-MNf2LjUt8dNQuYA_lTE)

If an administrator of your CryoSPARC instance has enabled access for you to modify the priority of a job (more on that below), you can set the priority when you queue a job via the Queue Modal. The priority of a job must be between 0-100.

![Set the priority of a job in the Queue Modal.](/files/-MNf2R-edCHY-ivb5cFe)

### Access & Default Priorities

When you create a new job in CryoSPARC, the priority value is automatically set using the default user and instance priority values. The default user job priority will take precedence; if it is not set, the default instance priority is used instead. If neither are set, the default priority is 0. You can manage both default priority values and whether a user can set them or not in the Admin Panel.

![Access the admin panel by clicking on your user account in the bottom right corner of the interface.](/files/-MNf2V0PHd5XBYs67k4u)

![Manage user access to modify the job priority via the "Manage Users" table](/files/-MNf2_2coU-O9ZWiqv_V)

#### Manage Access to Modify Job Priority

In the admin panel, administrators can enable and disable users from being able to modify the priority value of a job using the "Job Priority Management" column. Click on the **"Able to modify"/"Unable to modify"** button to change this value. Enabling this will allow users to see the "Job Priority" input box in the Queue Modal.

![Click on the "Unable to modify" button to enable modification of the job priority.](/files/-MNf2cowRaD6f3M_USP7)

#### Default User Priorities

You can set the default user priority by entering a value in the "User Default Job Priority" column. This value will be used by default whenever the user creates a job.

#### Default Instance Priorities

You can set the default instance priority by entering a value in the "Update Default Job Priority" input box. This will update the default job priority for all users in the instance. This value will be used if a user doesn't have their own default job priority set.

![Update the default job priority for all users in the instance.](/files/-MNf2nvDeHu4rgjzBFTA)

### Inspect Priorities of Your Jobs

You can view the priority of your jobs in several places. The easiest place to see it is in the Resource Manager, under the "Current Jobs" tab.

![The priority of a job is displayed in green next to an exclamation mark icon.](/files/-MNf301V4qBLA08DR5ZL)

You can also find the priority of jobs that previously ran in the "Job History" tab, under the "Priority" column.

![Click on a job to view its job card.](/files/-MNf33QV4NSB8T8-Fv5i)

When you select a job, you can also view its priority by looking at the Job Details panel on the right.

![The priority of this job is listed next to "Job Priority".](/files/-MNf382Y7nJotiQ-bQ86)

The final place you can find the priority of a job is in the job card's "Metadata" tab.

![Search for the "priority" key (CTRL+f/CMD+f) to find the job's priority value.](/files/-MNf3CV0-w7KHEZUcDJR)

## CryoSPARC Live Session Priority

CryoSPARC Live sessions also leverage the `priority` value. Any jobs created by the Live Session (CryoSPARC Live GPU Worker jobs, Streaming 2D Classification jobs, Ab-Initio Reconstruction jobs, Streaming Refinement jobs, etc.) will inherit the priority that was set for the Session.

By default, when you create a Session, the default priority set will follow the same method as when you create a new job: first, the user default will be used, then the instance default if a user default doesn't exist.

You can change the priority of a session by changing the "Priority" value in the Configuration tab in CryoSPARC Live.

![Set a Live Session's priority in the configuration tab.](/files/-MNf3KGERad_frdaPH3p)


# Guide: Configuring Custom Variables for Cluster Job Submission Scripts

{% hint style="warning" %}
The information in this section applies to CryoSPARC v4.1+.
{% endhint %}

User-set variables can be injected into cluster submission scripts before a job is queued. This may be useful for running jobs with different submission parameters on-the-fly, such as adjusting memory requirements in accordance with job resource demands, or keeping track of user quotas.

Custom variables can be configured at the instance level, the target level, and the job level.

Instance-wide and per-target level custom variables for cluster job submissions can be configured in the "Cluster Configuration" tab of the [Admin Panel](/application-guide/admin-panel).

<figure><img src="/files/vss3YLtWWvmyEB1FxYZk" alt=""><figcaption><p>The Cluster Configuration tab of the Admin Panel</p></figcaption></figure>

At the instance level, custom variables and their values will be applied to all targets which include the respective custom variables in their templates. If a target's submission script does not contain an instance level custom variable, that custom variable will not apply to the target.

At the target level, template variable names are parsed from the cluster submission script. Variables parsed from a target's script can be assigned values specific to that target. Any values for custom variables applied at the target level will override those that are set at the instance level. For example, if the custom instance level variable `var` is set to 'x' and the custom target level variable `var` is set to 'y', the value 'y' will be applied when the job is queued to the cluster.

At the job level, the values of instance level and target level custom variables can be modified before the job is queued. Modifications of the custom variable values at this level will override the default values set at the target level and instance level. Similarly to the previous example, if there is an instance level custom variable `var` set to 'x', there is a target level custom variable `var` set to 'y', and `var` is set to 'z' at the job level, then the value 'z' will be applied when the job is queued to the cluster.

To edit a custom variable at the job level, a dropdown menu will appear after selecting a cluster lane while queueing a job. Instance level and target level custom variables will appear here, as well as editable variables provided by CryoSPARC.

![](/files/wGaaUCJ0UhUVEyzhenhS)

## Reserved variable names

Below is a list of internal variables set by CryoSPARC that cannot be used as custom variables or be overridden by custom variables.

```
{{ run_cmd }}            - the complete command string to run the job
{{ num_cpu }}            - the number of CPUs needed
{{ num_gpu }}            - the number of GPUs needed. 
                           Note: the code will use this many GPUs starting from dev id 0
                                 the cluster scheduler or this script have the responsibility
                                 of setting CUDA_VISIBLE_DEVICES so that the job code ends up
                                 using the correct cluster-allocated GPUs.
{{ ram_gb }}             - the amount of RAM needed in GB
{{ job_dir_abs }}        - absolute path to the job directory
{{ project_dir_abs }}    - absolute path to the project dir
{{ job_log_path_abs }}   - absolute path to the log file for the job
{{ worker_bin_path }}    - absolute path to the cryosparc worker command
{{ run_args }}           - arguments to be passed to cryosparcw run
{{ project_uid }}        - uid of the project
{{ job_uid }}            - uid of the job
{{ job_creator }}        - name of the user that created the job (may contain spaces)
{{ cryosparc_username }} - CryoSPARC username of the user that created the job (usually an email)
{{ job_type }}           - CryoSPARC job type
{{ command }}            - used internally by CryoSPARC
```

## Modifying requested cluster resources with custom variables

`num_cpu`, `num_gpu`, and `ram_gb` are script variables provided by CryoSPARC specifying the job's estimated required resources needed to run. Sometimes before submitting jobs to a cluster, these values may need to be adjusted depending on the accuracy of the job's required resources estimation. Custom cluster script variables can be used in this case to modify the values requested of the cluster.

For example, below is a SLURM cluster configuration where the value for the amount of RAM required is modified by a `ram_multiplier` custom variable.

{% hint style="info" %}
Because custom variable values are injected in to the template as *strings*, one may have to apply a filter, like `|float`, if a numeric interpretation of the variable's value is required, as is the case in the `ram_multiplier` example below.
{% endhint %}

```
#!/usr/bin/env bash
#SBATCH --chdir={{ job_dir_abs }}
#SBATCH --export=NONE
#SBATCH --job-name cryosparc_{{ project_uid }}_{{ job_uid }}
#SBATCH --cpus-per-task={{ num_cpu }}
#SBATCH --gres=gpu:{{ num_gpu }}

## Example of a modifier variable
#SBATCH --mem={{ (ram_multiplier | default(1) | float * ram_gb) | int }}G

{{ run_cmd }}
```


# Guide: SSD Particle Caching in CryoSPARC

Overview of how SSD particle caching works, how much SSD space you need, configuration options and troubleshooting.

## Why is particle caching effective?

For classification, refinement, and reconstruction jobs that deal with particles, having local SSDs on worker nodes can significantly speed up computation: Many cryo-EM algorithms rely on random-access patterns and multiple passes though the data, rather than sequentially reading the data once. When you install CryoSPARC, you have the option of adding an `ssd_path`, which is a fast drive location on the worker node that particles will be copied to and read from when being processed. CryoSPARC manages the SSD cache on each worker node transparently.

When you run jobs that have the `Cache particle images on SSD` option turned on, particles will be automatically copied to and read from the SSD path specified. Furthermore, if multiple jobs within the same project require the same particles, the cache will be re-used and the copying step is skipped. If more space is needed, previously cached data will be automatically deleted. Setting up an SSD cache is optional on a per-worker node basis, but it is highly recommended. Nodes reserved for pre-processing (motion correction, CTF estimation, particle picking, etc.) do not need to have an SSD.

![](/files/-MNeHcnh-oydXMAbSwjg)

## Hardware

The size of your typical cryo-EM single particle datasets will inform the size of SSD you choose to use. To store the largest of particle stacks, we recommend 2TB SSDs. You can calculate the exact size of a particle dataset with the following calculation:

$$
Dataset\ Size =(4\*(box\_size^2)+nsymbt+header\_length)\*num\_particles
$$

For example: A 1,000,000 particle dataset with box size 256 will have a total size of 263.3 GB

$$
(4\*(256^2)+128+1024)\*1,000,000=263,296,000,000 \ bytes
$$

For example: A 2,000,000 particle dataset with box size 432 will have a total size of 1.5 TB

$$
(4\*(432^2)+128+1024)\*2,000,000=1,495,296,000,000 \ bytes
$$

## Configuration

### Installation

When installing CryoSPARC, you can use the parameter `--ssdpath` to specify the path of your SSD drive when you connect your worker to your instance. If you don't want to configure an SSD cache for a workstation node, specify the `--nossd` option.

```
bin/cryosparcw connect 
  --worker <worker_hostname> 
  --master <master_hostname> 
  --port <port_num>   
  --ssdpath <ssd_path>             : path to directory on local ssd
```

By default, if you specify the SSD path then the cache will be enabled with no quota or reserve.

### Advanced Parameters

You can specify two advanced parameters to fine-tune your SSD cache:

`--ssdquota`: The maximum amount of space that CryoSPARC can use on the SSD (MB)

`--ssdreserve`: The minimum amount of free space to leave on the SSD (MB)

The above options are useful when you're setting up CryoSPARC on a common compute node that will share the SSD with other applications.

### Updating Configuration

You can always update the SSD configuration at any time by running the `connect` command with the `--update` flag:

```
bin/cryosparcw connect
  --worker <worker_hostname>
  --master <master_hostname>
  --port <port_num>
  --update                         : update an existing worker configuration
  [--nossd]                        : connect worker with no SSD
  [--ssdpath <ssd_path> ]          : path to directory on local ssd
  [--ssdquota <ssd_quota_mb> ]     : quota of how much SSD space to use (MB)
  [--ssdreserve <ssd_reserve_mb> ] : minimum free space to leave on SSD (MB)

```

## Use

### Use the caching system when running a job

When you are running jobs that process particles (for example: Ab-Initio, Homogeneous Refinement, 2D Classification, 3D Variability), you will find a parameter at the bottom of the job builder under "Compute Settings" called `Cache particle images on SSD`. Turn this option off to load raw data from their original location instead.

![](/files/-MNeHjnk6-x7NHEbFGp9)

### Set a default parameter for the project

By default, the `Cache particle images on SSD` parameter is always on for every job you build, but if you'd like to keep this option off across all jobs in a project, you can set a project-level default.

In v2.15+, the parameter can be adjusted from the sidebar when a project is selected.

![](/files/-MNeHmE0qcNhQO971UfV)

In earlier versions of CryoSPARC, you can adjust this parameter by running the following command in a shell on the master node:

`cryosparcm cli "set_project_param_default('PX', 'compute_use_ssd', False)"`

where `'PX'` is the Project UID you'd like to set the default for (e.g., `'P2'`)

You can undo this setting by running:

`cryosparcm cli "unset_project_param_default('PX', 'compute_use_ssd')"`<br>

## Tips and Tricks

### Consolidating a Particle Stack

When caching a particle stack that is larger than space available on your SSD, you may optionally consolidate your particle stack. This option works if the current particle stack is a subset of the original particle stack. For example, when the cache reports how much data it's requesting to copy (`SSD cache : cache requires 1000000.00 MB more on the SSD for files to be downloaded.` & `SSD cache : cache successfully requested to check 2000000 files.`) and the sizes it reports seem much larger than you expected, you can consolidate your particle stack such that only the particle subset you care about is cached.

You might run into this situation if you ran an "Inspect Picks" job after an "Extract From Micrographs" job, and you modified the picking thresholds of your particles to include a smaller subset than the original stack.

You might also run into this situation after a round of 2D Classification. When you select classes, you create metadata that specifies which subset of the particle stack to use. When using this particle subset in further processing, the caching system will require the entire stack of particles to be cached, even though only the smaller subset is required.

To consolidate your particle stack, build a "Downsample Particles" job, connect your particles, and run the job. There is no need to change any parameters - nothing will change about your particle dataset except for the `.cs` metafile that will be recreated to reflect the smaller subset. You can use this smaller dataset to continue processing.

### Dynamic SSD Cache Paths

On some systems it is not possible to know the SSD cache path ahead of time. Instead, a dynamically-generated path is available for jobs to use at run-time.

To prompt a job to use this path, make the path available via a system-defined environment variable. Open `cryosparc_worker/config.sh` for editing and set the value of `CRYOSPARC_SSD_PATH` in the worker environment config to this variable:

```
# cryosparc_worker/config.sh
export CRYOSPARC_LICENSE_ID="xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
export CRYOSPARC_USE_GPU=true
export CRYOSPARC_CUDA_PATH="/usr/local/cuda"
export CRYOSPARC_SSD_PATH="$CUSTOM_DYNAMIC_SSD_PATH"
```

### Increase or Reduce Cache Files Lifetime

As of CryoSPARC v3.3, jobs that require cache automatically remove cache files that have not been accessed in over 30 days. If projects on your instance get actively worked on for more or less time, you may change the cache file lifetime by adding the following line to `cryosparc_master/config.sh`:

```
export CRYOSPARC_SSD_CACHE_LIFETIME_DAYS=15
```

Substitute `15` with the number of days your projects typically get worked on.

### Leveraging Multiple Threads to Copy Particles

In CryoSPARC v4.3.0+, multiple threads (default 2) are used to copy particles from the project directory to the local cache device. To modify the number of threads used, add the following line to `cryosparc_worker/config.sh,` where `num_threads` is the number of threads (e.g., 12) to spawn to copy files:

```
export CRYOSPARC_CACHE_NUM_THREADS=num_threads
```

Specify `export CRYOSPARC_CACHE_NUM_THREADS=1` to turn off multithreading and copy particle files sequentially in the main process.

## Troubleshooting

### `SSD cache : cache waiting for requested files to become unlocked.`

This temporary message usually means the files this job is trying to access are currently being cached by another job. For example, if you started two different Refinement jobs at the same time on the same node (Job A and Job B) using the same particle stack that haven't been cached on SSD before, both jobs try to first copy all particles onto the SSD. If Job A acquires the lock for the files first, it starts copying them and Job B shows this message. When Job A finishes copying the files, it unlocks them. Job B is unlocked and finds that the particles are already on the SSD, so it skips over the copy step.

### `SSD cache : cache does not have enough space for download... but there are no files that can be deleted.`

This message means that there is another CryoSPARC job or another application on the workstation taking up space on the SSD. If the former, the job showing this message will try to free up space as soon as it can, and it will continue processing. If there are files on the SSD that are not owned by CryoSPARC, it will not be able to delete them. It may be necessary to delete them manually.

## FAQ

**Is it safe to manually delete cache files for completed or unqueued/cleared jobs? Also, can I pre-cache with symlinks to skip caching?**

Yes, it is safe to delete cache files any time (it’s a read-only cache) and yes, the cache checks to see if files exist just based on path/size/modification date so symlinks should cause it to skip. Though it may be easier to just set the SSD Cache parameter to False in each job that you queue up.

Source:

[How to clear the cache in v2?](https://discuss.cryosparc.com/t/how-to-clear-the-cache-in-v2/2161)


# Guide: Data Management in CryoSPARC (v4.0+)

An overview of all data management utilities and common use cases.

{% hint style="warning" %}
The information in this guide applies to CryoSPARC v4.0+. For information about managing project directories in older versions, see [Guide: Data Management in CryoSPARC (≤v3.3)](/guides-for-v3/tutorial-data-management-in-cryosparc)
{% endhint %}

{% hint style="danger" %}
Do not remove from the filesystem any directory that is managed by an[ attached CryoSPARC project](#2.-attaching-detaching-archiving-and-unarchiving-projects). First, either

* perform the *Detach Project* CryoSPARC GUI action for unwanted project(s)
* or, to delete an unwanted project from the CryoSPARC GUI **and erase the project's data from disk**, perform the *Delete Project* GUI action
  {% endhint %}

{% hint style="info" %}
For additional data cleanup utilities available in v4.3+, please see: [Guide: Data Cleanup (v4.3+)](/setup-configuration-and-management/software-system-guides/guide-data-cleanup-v4.3)
{% endhint %}

## Overview

Single particle cryo-EM projects and labs continue to operate at increasing scales. In CryoSPARC v4.0, we introduce improved workflows and tools for dealing with archiving, transferring, exporting, and importing projects. Tools introduced in previous versions of CryoSPARC for exporting and importing individual jobs and individual results of various types (particle stacks, exposure stacks, volumes, etc.) remain available.

In CryoSPARC v4.0, the most important changes that have been made are:

* CryoSPARC projects are now explicitly **locked** (or **attached**) to a single CryoSPARC instance at a time. In previous versions, it was possible for a project directory to be imported into and accidentally modified by two instances at the same time, causing metadata corruption. Now, each project directory contains a lock file that marks the projects as in-use by (i.e., **attached to**) a particular instance.
* CryoSPARC project directories are now named based on the project title at creation time, rather than a numeric project-UID (e.g., "P12"). Numeric UIDs are only used to refer to each project within a single instance. This serves to ensure that project directories are user-recognizable, and that the numeric UID is not retained when a project moves from one instance to another.
* The life-cycle of a CryoSPARC project is now separated from the CryoSPARC instance(s) that interact with the project. The following lifecycle actions can be taken on a project:
  * **Create:** an instance can create a new project, and that project lives in a unique and self-contained project directory on disk. The project directory is created at creation time of the project. The project is attached to the instance that created it (and therefore there is a lock file present in the project directory). Within the instance to which the project is attached, the project has a unique numeric UID.
  * **Detach:** a user can opt to detach a project from the instance to which it is currently attached. This action ensures that no jobs or background processes are running in the project, and then removes the lock file from the project directory. In the UI, the project that was previously attached displays as "Detached" and can no longer be interacted with.
  * **Attach**: a project directory that has previously been detached (and therefore has no lock file present) can be attached to an instance. When attachment is performed, all the workspaces, jobs, and Live sessions within the project directory are imported into the attaching instance, and the project becomes usable within this instance and is given a new numeric UID. A lock file is written to the project directory. Attach takes the place of the previous **Import** action.
  * **Archive:** A project that is attached to an instance can be "archived" without detaching the project. This instructs CryoSPARC that the project directory is no longer available for reading and writing at it's current location, but will become available (possibly at a new location) at some future time. A user should archive a project before moving the project directory to a different location on disk, for example a different filesystem, a backup location, tape archive, or cold-storage. Archiving ensures that no jobs or background processes are running in the project and marks the project as archived, but does not remove the lock file from the project. Once archived, the project can still be browsed in the UI, but can not be modified.
  * **Unarchive:** A project that has been archived can be resurrected in the same instance from which it was archived. When unarchiving, the user is prompted to provide the (possibly changed) location of the project directory on disk. For example, a project can be archived and then the project directory moved to a cold unaccessible backup. Later, when needed, the project directory can be restored to an accessible filesystem, and the project can be unarchived pointing at the new project directory location. This makes the project available once again for further processing.
* As a minor change, output files of CryoSPARC jobs that are stored in job directories no longer have the `cryosparc_PXX_` prefix, since the numeric project UID can change when a project moves from one instance to another. In order to retain the existing behaviour of the CryoSPARC UI and limit confusion between different files, when CryoSPARC results are downloaded through the browser in the UI, the prefix `cryosparc_PXX_` is added to the local filename of the download in the browser, using the then-current numeric project UID.

These changes make it much simpler to perform the following actions:

* **Detaching** a project from one CryoSPARC instance and **attaching** it to another instance
* **Sending** a project initially started at a centralized facility to a user who is going to continue processing at home in their own instance, by detaching and attaching
* **Archiving** a project to remote/slow storage for later retrieval and resurrection
* **Copying** a project directory to make a complete clone
* **Changing** the name or location of a project directory on disk, by archiving and unarchiving

The following actions remain possible, with no change in behaviour in v4.0:

* **Reducing** the disk space used by a project by removing intermediate results created by jobs
* **Uploading** the final results of a CryoSPARC job to an online repository
* **Advanced manipulation** of CryoSPARC metadata at a low level or programatically, by exporting a results (e.g., a particle stack), manually modifying the associated .cs files, and importing again

The following sections describe specific aspects of data management in CryoSPARC in more detail.

## 1. CryoSPARC Projects and Project Directories

### Projects and "Continuous Export"

CryoSPARC workflows are naturally divided into projects. Each project should contain the work and jobs for one or more related data collection sessions that are associated with a given sample/target. Project boundaries are strict, in the sense that files and results from one project cannot be directly used in another project. Project directories are self-contained, and all image processing data (except for imported raw data, see [#7.-imported-data-in-project-directories](#7.-imported-data-in-project-directories "mention")) pertaining to a project is written to the project directory.

A project directory always contains all the information needed to define that project. The project directory is written to every time certain actions are taken within CryoSPARC, for example changing project, workspace, or job metadata (titles, descriptions, etc), and when jobs complete processing. This **"continuous export"** model ensures that at any time, a project directory is self-contained and if anything goes wrong with a CryoSPARC instance or database, the projects remain intact and up-to-date, without the user having to manually trigger an export action.

As a safety feature, a project directory can, at any time, be transferred/renamed/copied and **attached** as a valid project in any (other) CryoSPARC instance that can read the files. **This is true even if the original CryoSPARC instance that created the project is no longer functional.** See [#use-case-rescuing-a-project-from-an-inoperable-instance](#use-case-rescuing-a-project-from-an-inoperable-instance "mention") for details on how to rescue a project from a failed or inoperable instance.

Similar to projects, jobs inside the project are stored in a self-contained format, and job directories are updated whenever the jobs are created, modified, or completed. **Note:** jobs that are in `launched`, `started`, `running`, `waiting`, `killed` or `failed` status will not be updated on disk until they enter `completed` status - either by actually completing, or by the user choosing the `mark as completed` option in the Job Details panel.

### **Project directories and lock files**

For projects created in CryoSPARC v4.0+, the project directory is initially named based on the title of the project at creation time. For example, if a project is titled\
"My Protein, Data Collection (October 1 2022)"\
the project directory will be created as\
`CS-my-protein-data-collection-october-1-2022`.\
The project directory will be created within the container directory indicated at creation time. The project directory can be changed later on (see [#use-case-renaming-a-project-directory](#use-case-renaming-a-project-directory "mention")).

Inside each project directory in CryoSPARC v4.0+ (including existing projects), there will be a lock file present called `cs.lock`. **This file should not be removed or changed.**

###

## 2. Attaching, Detaching, Archiving, and Unarchiving Projects

### **Attach**

Detached projects can be **attached** to a CryoSPARC instance. Attaching a project creates a new project in the instance using the indicated project directory. All project details, workspaces, jobs, and sessions in the detached project will be made accessible, and the project directory will be treated as an active project directory by the instance. Any intact CryoSPARC project directory that does not contain a lock file (including from a previous CryoSPARC version) can be attached.

{% hint style="warning" %}
Note: Projects already belonging to another CryoSPARC instance **cannot** be attached until they are detached from their original instance. When this is not possible see [#use-case-rescuing-a-project-from-an-inoperable-instance](#use-case-rescuing-a-project-from-an-inoperable-instance "mention")
{% endhint %}

Projects can be attached under the “New Project” dropdown menu:

<figure><img src="/files/WH3UWhDpXG91zd0NlEgQ" alt=""><figcaption></figcaption></figure>

Once the attach process begins, you will see a new project appear in the projects page. This project will have a new numeric UID (distinct from the numeric UID of the project in the instance where the project previously was attached). Once attachment is complete, users can begin to interact with the project and continue processing.

### Detach

Projects can be **detached** from their CryoSPARC instance. Detaching a project unlocks the project from its instance, allowing the project folder to be moved to another location or attached to another instance. In the UI of the instance where the project is being detached, the project will also display as ‘detached’ and no longer be usable. A detached project’s details, workspaces, jobs, and sessions are saved to the project directory.

A project can be detached using the “Detach Project” button in the project’s “Actions” menu:

<figure><img src="/files/9SyAwxteFIGaqClSL7pr" alt=""><figcaption></figcaption></figure>

Detached projects will show an icon on their cards:

<figure><img src="/files/t9uEgf2qKsasXJjmTeMQ" alt=""><figcaption></figcaption></figure>

Upon detaching a project, the project is no longer associated with the CryoSPARC instance, but some project information is retained in the CryoSPARC database. As of v4.1.2, the “Delete Project from Database” action can be used to remove the remaining database entries associated with the project. Performing this action on a detached project hides it in the UI and removes large database files, potentially freeing up space on disk.

### Archive

{% hint style="warning" %}
This section describes the CryoSPARC *Archive Project* function. *Archive Project* does **not** copy the project data. Ensure that *Archive Project* is followed by copy/transfer of the project directory to long-term storage *outside* CryoSPARC, as needed.
{% endhint %}

{% hint style="warning" %}
One must not change the contents of a CryoSPARC project directory outside CryoSPARC. *Unarchiving* a project whose project directory has been modified since archiving will lead to inconsistencies between the project directory and the project's database records. Such inconsistencies can lead to CryoSPARC malfunction and data loss.
{% endhint %}

CryoSPARC projects can be **archived** to allow their project folder to be moved on disk and un-archived at a later date. Archiving sets the project to read-only mode, where it can be seen in the UI but cannot be modified. All project details, workspaces, jobs, and sessions will be maintained in the CryoSPARC database as well as in the project directory on disk. CryoSPARC does not expect the project directory to be available for read or write while a project is archived. When moving a project directory, be sure to consider moving the raw data that was imported into the project as well (see [#7.-imported-data-in-project-directories](#7.-imported-data-in-project-directories "mention"))

A project can be archived using the “Archive Project” button in the project’s “Actions” menu:

<figure><img src="/files/nUGT0ximPd43YAPqywGK" alt=""><figcaption></figcaption></figure>

Archived projects will show an icon on their card:

<figure><img src="/files/fIrdf75c0zAjhazUdNks" alt=""><figcaption></figcaption></figure>

The archived status can also be seen in the project details:

<figure><img src="/files/zWz9oDQFLDtNPhE0PgF0" alt=""><figcaption></figcaption></figure>

### Unarchive

Archived projects can be unarchived back into the CryoSPARC instance, removing the read-only status and allowing the project to be modified again. Projects can be unarchived from a different project directory location than the location at time of archive.

{% hint style="warning" %}
Note: Archiving and unarchiving should only be used with the intention of keeping the project tied to the current instance of CryoSPARC. Users looking to transfer projects between CryoSPARC instance should refer to the Attach and Detach features instead.
{% endhint %}

{% hint style="warning" %}
*Unarchiving* a project whose project directory has been modified since archiving will lead to inconsistencies between the project directory and the project's database records. Such inconsistencies can lead to CryoSPARC malfunction and data loss.
{% endhint %}

A project can be unarchived using the “Unarchive Project” button in the project’s “Actions” menu:

<figure><img src="/files/MqER2rdTB77xpqPAok1P" alt=""><figcaption></figcaption></figure>

When unarchiving a project, the project directory must be specified:

<figure><img src="/files/TdRMNhby0bVnaMf6IErD" alt=""><figcaption></figcaption></figure>

## 3. Ability to view instance storage statistics

{% hint style="info" %}
This functionality has not changed in v4.0. See [/pages/-MNiDCV\_2xPWJWCTbMK7#3.-ability-to-view-instance-storage-statistics](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/pages/-MNiDCV_2xPWJWCTbMK7#3.-ability-to-view-instance-storage-statistics "mention")
{% endhint %}

## 4. Ability to clear intermediate results

{% hint style="info" %}
This functionality has not changed in v4.0. See[/pages/-MNiDCV\_2xPWJWCTbMK7#4.-ability-to-clear-intermediate-results](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/pages/-MNiDCV_2xPWJWCTbMK7#4.-ability-to-clear-intermediate-results "mention")
{% endhint %}

Several job types (2D Classification 3D Classification, and 3D Variability Analysis) have an option to control whether the job will save intermediate results at all. By default, jobs will save intermediate results. However, this can be turned off on a per-job level using job parameters, or it can be turned off at the project level for each job type. To do so, select the project and at the bottom of the details panel, set job-specific defaults under the 'Generate Intermediate Results' module:

<figure><img src="/files/Dab8BHao8GKoyPNHR2ST" alt="" width="375"><figcaption><p>Project-level defaults for generation of intermediate results</p></figcaption></figure>

## 5. Ability to export and import individual jobs

{% hint style="info" %}
This functionality has not changed in v4.0. See [/pages/-MNiDCV\_2xPWJWCTbMK7#5.-ability-to-export-and-import-individual-jobs](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/pages/-MNiDCV_2xPWJWCTbMK7#5.-ability-to-export-and-import-individual-jobs "mention")
{% endhint %}

## 6. Ability to export and import low-level output groups

{% hint style="info" %}
This functionality has not changed in v4.0. See [/pages/-MNiDCV\_2xPWJWCTbMK7#6.-ability-to-export-and-import-low-level-output-groups](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/pages/-MNiDCV_2xPWJWCTbMK7#6.-ability-to-export-and-import-low-level-output-groups "mention")
{% endhint %}

## 7. Imported data and symlinks in project directories

When raw data is imported into a CryoSPARC project (using an `Import` Job), the raw data is not copied into the project directory. Rather, symlinks are created within the Import Job directory pointing to the raw data. Aside from these symlinks, CryoSPARC jobs do not create symlinks that point to locations that are outside the project directory. This keeps project directories self-contained.

The symlinks within import jobs can be changed if the position of the raw data on disk changes. For example, when Archiving a project, if the raw data (e.g., raw movies) are also archived to a different location than where they were imported from, the import symlinks must be updated. See the guide here for more details:[/pages/-MNeMxOeXMsIWu8jTuYb#a.-moving-only-raw-particle-micrograph-or-movie-data-already-imported-into-cryosparc](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/pages/-MNeMxOeXMsIWu8jTuYb#a.-moving-only-raw-particle-micrograph-or-movie-data-already-imported-into-cryosparc "mention")

Due to the use of symlinks, it is important that when copying or moving a project directory, symlinks **NOT** be dereferenced (i.e., do not use the `-h` flag with `tar` ). If symlinks are dereferenced, the new copy of the project directory will also contain copies of all the raw data files, as well as potentially multiple copies of intermediate and output files that are internally symlinked within the project directory. Instead of dereferencing symlinks, raw data should be archived separately from project directories.

## Use Cases and Examples

### Use Case: Moving a project directory from one storage location to another

Sometimes, you may need to move a project directory on disk. You may have created it in the wrong place accidentally, you may have a full disk, or if you have tiered storage, e.g., a fast SSD-backed storage system for active projects and a slower HDD-backed storage array for bulk storage, you may wish to move a project directory from the fast filesystem to the slower filesystem once most processing is complete.

In these cases, you can simply:

1. Archive the project
2. Move the project directory to its new location
3. Unarchive the project using the path to the project directory at its new location

The project will now be usable once again, and all reads/writes will happen to the new project directory location.

### Use Case: Renaming a project directory

In CryoSPARC v4.0+, project directories are named based on the project title entered at creation time. If you later change the title, you can rename the project directory using the following steps:

1. Archive the project
2. Rename the project directory on disk, but leave it in it's original location
3. Unarchive the project using the path to the project directory with its new name

### Use Case: Transfer a project from one CryoSPARC instance to another

When you need to move a CryoSPARC project between instances, for example when transferring a project from a data collection facility to a user's home facility, use the following steps:

1. Detach the project from its original CryoSPARC instance
2. Copy the project directory to a location accessible by the new CryoSPARC instance
3. Attach the project to the new CryoSPARC instance using the path to the project directory at its new location
4. (Optional, available in v4.1.2+) Use the “Delete Project from Database” action on the detached project to remove and remaining database entries relating to this project from its original CryoSPARC instance

### Use Case: Archive or Detach a project directory and consolidate it for long term storage

Once processing in a project is complete, the project can be either detached (if it is unlikely to be brought back to the same CryoSPARC instance) or archived. Either action will allow the project directory to be moved or compressed without causing errors in the CryoSPARC instance.

Be sure to separately archive/copy/move the raw data that was imported into the CryoSPARC project, as raw data is not stored inside the project directory. See [#7.-imported-data-and-symlinks-in-project-directories](#7.-imported-data-and-symlinks-in-project-directories "mention")

The project directory can be copied as-is, and stored on a backup, remote, or cold-storage filesystem. In some cases it may help to `tar` the project directory into a single file. An example command to consolidate a project directory is:

```
cd /u/cryosparcuser/cryosparc_projects/
tar -cvf P47.tar ./P47
```

Note that you can use any method you choose to archive/transfer/store the project directory, as long as the entire contents remain intact.

If you need to access the project at a later date, you can un-`tar` the bundle to any accessible filesystem. Then, if the project was archived (you can tell by checking that the `cs.lock` file is still present inside the project directory), you can un-archive it to the same instance from where it was archived. Otherwise if it was detached, you can attach it in any instance.

### Use Case: Rescuing a project from an inoperable instance

If a CryoSPARC v4.0+ instance is no longer operable (due to database corruption or other issue), a project that was attached to that instance can be rescued by attaching to a new instance. Use the following steps:

1. Ensure that that inoperable instance is completely shut down, and that there are no remaining "zombie" processes associated with that instance still running.
2. For additional safety, make a copy of the project directory to be rescued and use the copy for subsequent steps.
3. Delete the `cs.lock` file in the project directory.
4. In the new instance, use **Attach Project** and point to the project directory where the lock file was removed.
5. The new instance should import all available workspaces, jobs, and sessions and make the project directory available for use once again.

### Use Cases that are unchanged in v4.0

The following use cases remain unchanged in v4.0+:

* [Guide: Data Management in CryoSPARC (≤v3.3)](/guides-for-v3/tutorial-data-management-in-cryosparc#use-case-share-a-particular-job-with-another-user)
* [Guide: Data Management in CryoSPARC (≤v3.3)](/guides-for-v3/tutorial-data-management-in-cryosparc#use-case-upload-your-particle-stack-to-empiar)
* [Guide: Data Management in CryoSPARC (≤v3.3)](/guides-for-v3/tutorial-data-management-in-cryosparc#use-case-manually-modify-cryosparc-outputs-and-metadata-for-continued-experimentation)


# Guide: Data Cleanup (v4.3+)

New features in v4.3+ for managing and cleaning up project data.

## Overview

Single particle cryo-EM datasets can be large (multiple TB), and the additional data generated by processing can grow equally large. Managing project data and cleaning up at various points in the project life cycle are important aspects of successful cryo-EM operations.

CryoSPARC v4.3 introduces several new features for managing and cleaning up project data. This guide provides a conceptual overview of data generated during processing and outlines recommended strategies for several use cases.

## CryoSPARC project data, job data, and cleanup strategies

CryoSPARC creates a self-contained project directory each time a new project is created, and all files generated by CryoSPARC related to a project will be stored in its project directory. For details about the project life cycle, see:

{% content-ref url="/pages/F3KBgDxkuaoVRFwpV0KW" %}
[Guide: Data Management in CryoSPARC (v4.0+)](/setup-configuration-and-management/software-system-guides/guide-data-management-in-cryosparc-v4.0)
{% endcontent-ref %}

Similarly, each time a CryoSPARC job is created, a job directory is created inside the associated project and the job data is stored in that job directory. When a job is **cleared**, the job directory is emptied but the job’s metadata (parameters, inputs, etc) are kept in the CryoSPARC database. This allows cleared jobs (or chains of cleared jobs) to be re-run at a later time.

Like jobs, CryoSPARC Live Sessions also have a session directory that is created inside their associated project directory, and session data (micrographs, particles, etc) are stored in the session directory.

In the standard project life cycle, projects (including all their jobs and sessions) can be detached from an instance or archived if they are not needed for some time. However, this does not generally reduce the size of a project directory. The new tools described here allow for shrinking projects, workspaces, and Live sessions so they can be dealt with more efficiently.

These tools will delete project data and therefore must be used with care, but are designed to allow cleaning up project data with clarity and confidence about what exactly will get deleted.

{% hint style="info" %}
CryoSPARC does not currently (as of v4.3) attempt to manage *raw* data on your filesystems. That is, data imported into CryoSPARC using the import jobs is not copied into project directories and therefore CryoSPARC will not delete or modify it at its original location.
{% endhint %}

There are five ways that CryoSPARC can clean up project data. Each is summarized here and described in more detail below.

* **Clearing jobs that are not needed to achieve a final result:** One or more jobs in a project or workspace can be marked as **final**, meaning that that job and all ancestors providing input to that job should not be removed from the project. All other jobs (i.e., non-final jobs) can be cleared in one step to prune unnecessary branches in the processing workflow.
* **Compacting CryoSPARC Live Sessions:** Live sessions produce motion corrected micrographs and extracted particles. In v4.3, Live sessions can be **compacted** which removes all this data, and can be **restored** which uses saved parameters and particle locations to reprocess and reproduce the removed data at a later date. Before compacting, useful particles from a session can be retained separately by **restacking** the particles (see below).
* **Clearing preprocessing jobs:** Motion correction, CTF estimation and particle extraction jobs create large amounts of project data but can generally be safely cleared and re-run in the future (with the same parameters) in order to regenerate the data they produced. Before clearing particle extraction jobs, useful particles can be retained separately by **restacking** the particles (see below).
* **Clearing Intermediate Results:** Reduces the size of job directories by deleting unused intermediate data from them. This means removing results from early iterations of a job but keeping the results from the final iteration for downstream use.
* **Clearing killed and failed jobs:** Jobs that did not complete may still have produced sizeable outputs and can be cleared in one step.

## Project and workspace Cleanup Data tool

<figure><img src="/files/kWpcmiEpFVXLmIDzxztA" alt=""><figcaption></figcaption></figure>

The Cleanup Data tool can be used to clear project data in bulk using recommended strategies, and can be used at the project or workspace level. The options in the tool can be customized to match specific clearing preferences. **Details about each option can be found below.**

The Cleanup Data tool can be accessed at any time and is accessible via the quick actions menu or sidebar action panel when a project or workspace is selected.

{% hint style="info" %}
Whether performing a cleanup on a project or workspace, the available options are identical. The only difference is that the **workspace cleanup will only affect jobs that exist in the selected workspace**. Note that in v4.3, **linked jobs** (i.e., jobs that are in the workspace being cleaned, but are also linked in other workspaces) **will be treated as being part of the workspace being cleaned and therefore will be cleaned up by the tool, even if they appear in other workspaces as well.** For this reason, it is important to mark as final all the important results across a project **before** cleaning up each workspace. By doing so, ancestors of final results across the project will be preserved in every workspace.
{% endhint %}

The Cleanup Data tool is comprised of four components:

* The header provides an overview of what project or workspace will be affected, how many sessions (if any) are contained within the project/workspace, how many jobs are contained within the project/workspace and how much space they take up on disk.
* The left side of the dialog displays a checklist of all cleanup operations available, broken down into categories.
* The right side of the dialog displays a tabbed interface, the cleanup preview - each with insight into how the cleanup action will affect the project or workspace.
* The footer allows you to return to the browse interface or proceed with the cleanup action.

### Data cleanup preview

As you select options, the data cleanup preview will update providing an overview of how the action will affect the project or workspace size. The preview has multiple tabs:

<figure><img src="/files/trQWl6Cik9DrZrZrR9Lb" alt=""><figcaption><p>By default, none of the options will be selected. This allows for fine-tuning exactly what data should be cleared for a given project or workspace.</p></figcaption></figure>

* **Estimate:** a breakdown of which jobs will be cleared, which jobs will only have intermediate results cleared, which jobs will be deleted, and a total of all jobs that will be modified as a result of the cleanup action. Below is a visual bar cart representing the breakdown and a percentage estimate of how much the cleanup action will reduce the project directory size on disk.

<figure><img src="/files/5NIoXlwguYGCu7QQyU6I" alt=""><figcaption></figcaption></figure>

* **Preprocessing:** a list of preprocessing jobs, non-preprocessing jobs and total jobs in the project or workspace along with their size on disk and percentage makeup compared to the total size of all jobs. In v5.0+ The list of preprocessing jobs may be subdivided into 3 categories if the project or workspace contains any jobs marked as final: Non-final, Final Ancestor, and Final. This breakdown helps to show the size of preprocessing jobs that will not be cleared due to being marked as final or ancestor of a final job. Preprocessing jobs marked as final ancestor may be optionally cleared using the "Include final ancestor jobs" toggle.

<figure><img src="/files/SYh2omru01bbXNf26c7l" alt=""><figcaption></figcaption></figure>

* **Intermediate results:** a list of all jobs in the project or workspace (if any) that have produced intermediate results.

<figure><img src="/files/VtGNmEARzGEylaWGh34k" alt=""><figcaption></figcaption></figure>

* **Non-final jobs:** a breakdown of all jobs that are marked as final, ancestors of final and neither (non-final) along with their size on disk and percentage makeup compared to the total size of all jobs. See below for details.

<figure><img src="/files/aQvH3Ai66gU4WVZgFb1I" alt=""><figcaption></figcaption></figure>

* **Killed/Failed jobs:** a breakdown of jobs by status along with their size on disk and percentage makeup compared to the total size of all jobs.

{% hint style="danger" %}
The Cleanup Data tool **does not** modify jobs or data created by Live sessions. For reducing sizes of sessions on disk for cleanup or archival purposes, view the section below on [Live session compaction and restoration](#live-session-compaction-and-restoration)*.* Within the Cleanup Data tool, the number of sessions and their size on disk will be reported in the header section for reference.
{% endhint %}

### Data sources and refreshing size

The Cleanup Data tool uses two sources of information to provide insight into the project or workspace breakdown: the size of the project on disk and the size of individual jobs or sessions.

The project sizes are calculated automatically when a project is imported and is manually refreshed via the sidebar. Job sizes are kept track of automatically by CryoSPARC. When project size is out of date, you will see a notification bar under the header indicating that the size should be refreshed, and a button to do so.

{% hint style="danger" %}
**Migration from older instances:** when you upgrade from a previous version to v4.3, all projects will not have the required size statistics calculated. Therefore when you open the cleanup data tool on an older project, you will have to recalculate the project size before continuing.
{% endhint %}

### Clearing vs. deleting jobs

The Cleanup Data tool has checkboxes for selecting whether to only clear, or clear and delete jobs in each category that can be cleaned up.

**Clearing** a completed job is a straightforward way of deleting job data while keeping the job inputs, parameters, and metadata intact. A cleared job takes up relatively little space, and can be re-run if its results are needed in the future. Once a job is cleared, all of its output data will be deleted, and its status will be set to **building**. Child jobs that use a cleared job’s output data as input will no longer be able to run, as the necessary files for them to run are deleted. However, outputs created by child jobs that were previously run will not be affected. Re-running the cleared job and regenerating its outputs will allow child jobs to be run again.

Be aware that re-running a cleared job does not guarantee bit-for-bit identical outputs to those generated the first time. Many factors outside of a job’s parameters can influence its outputs. See [Determinism in CryoSPARC](#determinism-in-cryosparc) for more details.

### Data management info in sidebar view

For easier auditing of projects and workspaces, the data management module - visible when a single project or workspace is selected - displays a similar breakdown of jobs that are down in the preview view of the data cleanup action.

<figure><img src="/files/r9E27exwW8vQk3tZjTcW" alt=""><figcaption></figcaption></figure>

In the case where a project also contains sessions, a project directory breakdown will display how much of the project size on disk pertains to jobs, how much to sessions and other files (metadata or files transferred by a user that aren’t outputted by CryoSPARC):

<figure><img src="/files/9LUD9qYkLPqlRbOU0w46" alt=""><figcaption></figcaption></figure>

### 1. Clear preprocessing jobs

Clearing deterministic jobs, such as pre-processing and extraction jobs, can remove a lot of reproducible data and save significant space if their results are no longer needed for processing. When the relevant checkboxes are selected, the Cleanup Data tool will clear preprocessing jobs as long as they are not marked as a final result. In v5.0+ preprocessing jobs that are marked as an ancestor of a final result are likewise not cleared unless the "Include final ancestor jobs" toggle is toggled on. Before clearing extraction jobs (which produce particle data), it can often be helpful to **restack** useful particles (see below).

#### Determinism in CryoSPARC

Generating results deterministically is an important concept within scientific computing. Unfortunately achieving perfect determinism in high-performance applications is not trivial. While using CryoSPARC, many factors, including floating point precision, dependency versions, GPU models, and non-deterministic parallel execution can result in differences in outputs of jobs with identical inputs.

Preprocessing jobs (motion correction, CTF estimation, particle picking, extraction) can generally be considered deterministic in the sense that, **for the same CryoSPARC version**, clearing a job and re-running it later with the same input data and the same parameters will yield output data that is nearly identical, with the differences (for the reasons mentioned above) smaller than the measurable signal in cryo-EM data.

Other downstream job types, for example 2D classification, 3D ab-initio reconstruction, 3D refinements, 3D variability analysis, etc. can generally not be considered deterministic. This is because 1) these jobs heavily use stochastic algorithms that depend on random numbers and so forgetting to or incorrectly setting a random initial seed will produce very different results, and 2) because even with the same random seed, these jobs are all iterative and so small unavoidable differences (due to the reasons mentioned above) will accumulate and amplify over iterations leading to measurably different final results.

For these reasons, the Cleanup Data tool allows to clear preprocessing jobs that can be safely re-run in the future but does not attempt to clear downstream jobs.

#### Restack particles

Extracted particle sets take up significant space on disk, and the particles are arranged into files with typically one file per micrograph. Over the course of processing, only a subset of initially extracted particles are typically used for downstream processing. Particles are filtered out during 2D and 3D classification and sometimes particles from entire micrographs may be discarded during curation. Filtering out particles unfortunately does not erase them from disk, and a final particle set will typically be sparsely spread out through all of the initially extracted files.

At any point in a project when a useful subset of particles has been identified, the particles can be **restacked** using using the [Restack Particles](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/extraction/job-restack-particles) job. This job takes in an input particle set and writes new particle files containing only those particles as the output. After restacking, the original particle files can be deleted by clearing the extraction job from which they were produced. This can often reduce disk usage substantially and also has performance advantages for caching.

As an example of how use Restack Particles:

* A particle set is created as the output of Extract From Micrographs
* 2D Classification is run on the particle set, followed by Select 2D Classes, resulting in a subset of filtered particles of the original set as the output
* Restack Particles can now be run using the particle output of Select 2D Classes, producing new files with only those particles
* The original Extract From Micrographs job can be cleared (either manually, or automatically by the Cleanup Data tool) and processing can continue using the output of Restack Particles

### 2. Clear intermediate results

{% hint style="info" %}
By default, as of v4.3, automatic deletion of intermediate results is enabled for new jobs in all projects. This step is only applicable for users who have disabled the deletion of intermediate results at the project or individual job level.
{% endhint %}

The Cleanup Data tool will, with the appropriate checkbox checked, clear out intermediate results from all jobs in the project or workspace. This will not stop downstream jobs from being able to be created or run, since final results of jobs will not be cleared. Likewise, if an intermediate result was used in a downstream job, that particular intermediate result will not be cleared.

Separately, intermediate results can be manually cleared at the project, workspace, or job level using the button in the Actions menu.

For more details on clearing intermediate results, please see:

{% embed url="<https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/guide-data-management-in-cryosparc-v4.0+#4.-ability-to-clear-intermediate-results>" %}

### 3. Clear non-final results

Often during the course of processing data, especially in advanced stages, you may experiment with multiple different pathways or attempts with different jobs and parameter to explore the data. Often only one of these will be fruitful and many branches in the processing tree will be redundant.

CryoSPARC v4.3 makes it easy to clean up unnecessary processing branches. At any point during processing, when a significant result or goal is achieved by a job in a CryoSPARC project, that job can be marked as a “final result”. The job will show with a special flag indicator for **final result** and all ancestors of the job will also be automatically marked as **ancestor of final result.**

To mark a job as final, open the quick actions menu or sidebar actions panel for that job and select the "Mark job as final" option. Alternatively, the keyboard shortcut `Alt/Opt` + `Shift` + `F` can be used to mark a selected job as final.

<figure><img src="/files/L4uvNWBnc5gDBcc4V884" alt=""><figcaption><p>Right-click menu with option to mark job as final</p></figcaption></figure>

<figure><img src="/files/Pn5Mru8jM39Gy8DpCj8R" alt=""><figcaption><p>Job marked as final</p></figcaption></figure>

{% hint style="danger" %}
It is best to mark important jobs as final before performing data cleanup actions, to avoid any loss of important results.
{% endhint %}

These jobs are treated specially by the Cleanup Data tool:

* **Final results:** Jobs marked as a **final** result cannot be cleared, cannot have their parameters changed, and cannot be deleted. **Their data is protected from the Cleanup Data tool and from manual actions within CryoSPARC as well.** When a job is marked as final, it is marked as final in all workspaces where it appears.
* **Ancestor of final results:** Ancestors of jobs marked as a final result are given a special status, called **ancestor of final**. When a job is marked as final, its ancestor jobs will be marked across the entire project, regardless of whether they appear in the same or a different workspace from the final job.

At any point in a project lifecycle, you can mark important jobs as **final**, and the Cleanup Data tool can be used to clear or delete non-final jobs (meaning jobs that are not final and are not ancestors of final jobs). **When run at the project level, all non-final jobs in the project will be affected. When run at the workspace level, only non-final jobs in the workspace will be affected.** Running the Cleanup Data tool with the “Clear non-final jobs” checkmark checked will effectively prune the processing tree, keeping all jobs necessary to achieve the final results but no others.

{% hint style="info" %}
When the Cleanup Data tool is run with only the “Clear non-final jobs” checkboxes checked, it will preserve ancestors of final jobs, but will clear other unnecessary branches. However, when the Cleanup Data tool is run with any of the the “Clear pre-processing jobs” checkboxes checked, it may clear ancestors of final jobs that fall into those categories (e.g. motion correction, CTF estimation, etc). v5.0+ introduces a "Include final ancestor jobs" below the “Clear pre-processing jobs” checkboxes to control whether these jobs should be cleared or not.
{% endhint %}

#### Manually selecting chains and ancestors

In CryoSPARC v4.3, along with automatically using the Cleanup Data tool to trace jobs and their ancestors, it is also possible to manually select chains of jobs or ancestors/descendant of jobs and perform actions on those selections.

* **Select a chain of jobs that are connected to each other**
  * Select the last job in the chain, and then `cmd`/ `ctrl` click the first job in the chain you wish to select
  * Right click on either job and choose “Select Job Chain”. The selection will update to include all jobs in the chain between the first and last job.
  * Right click on any job in the chain and perform an action on all the jobs such as moving, linking to another workspace, cloning, clearing, deleting, etc.
* **Select ancestors or descendants of a job**
  * Select a job, right click and choose “Select ancestor jobs” or “Select descendant jobs”
  * The selection will be updated to include all ancestors or descendants
  * Right click on any job in the set and perform an action on all the jobs such as moving, linking to another workspace, cloning, clearing, deleting, etc.

For more information about multi-selecting jobs, see:

{% embed url="<https://guide.cryosparc.com/application-guide-v4.0+/managing-jobs#multi-actions>" %}

### 4. & 5. Clear killed and failed jobs

Killed and failed jobs may have generated data while they were running which can be cleared. These jobs may be in this status because of bad inputs or errors and usually have no usable outputs. The Cleanup Data tool contains options to clear these killed and failed jobs.

In some cases, for jobs that generate intermediate results, these intermediate results may be usable as inputs to other jobs (e.g. using intermediate iterations in a classification job). In such a case, the killed or failed job can be marked as completed by right clicking on the job card and selecting “Mark Job as Complete”. This will allow for its intermediate results to be used as inputs to other jobs, and prevent it from being cleared when clearing killed or failed jobs with the Cleanup Data tool.

## Live session compaction and restoration

Live sessions generate large amounts of project data in the form of motion corrected micrographs and particle stacks. This data is stored in the session directory within the project directory.

In CryoSPARC v4.3, there are now actions available to reduce the disk space used by Live sessions once a project has reached a stage where the preprocessing data and particle extraction does not need to be repeated. Live sessions that are compacted can be restored in case preprocessing or extraction does need to be repeated.

{% hint style="info" %}
If you wish to continue being able to perform downstream processing with a final particle stack that was initially extracted by CryoSPARC Live, it is best to **restack** the particles before compacting the Live session (see above).
{% endhint %}

### Compacting a Live session

Compacting can be done after marking a Live session as completed via the session’s Actions menu. Compacting will clear pre-processing and extraction stages, while saving extracted particle locations, all parameters, and user-inputted data (e.g. manual particle picks). A session cannot be modified after it is compacted, until after it is restored.

<figure><img src="/files/U09MjK7KmpI7LtP64aHo" alt=""><figcaption></figcaption></figure>

{% hint style="warning" %}
After compaction, the Live session will not be functional on CryoSPARC versions less than v4.3.0. This means that if CryoSPARC is downgraded or if the project is detached and reattached to an instance of a lower version, the session will not be able to be restored until it is brought back to a v4.3+ instance.
{% endhint %}

### Restoring a Live session

A compacted session can be restored via the session’s Actions menu. The restoration process brings a session back to it’s pre-compaction status by re-running pre-processing stages for all exposures, which can take significant time. Session restoration requires a lane to run on, which can be set via the Configuration tab of a session.

Should the restoration process be interrupted (e.g. by a failed worker job), it can be resumed by initiating the restoration process again. If individual exposures fail during the restoration process, “Reset failed exposures” can be used to retry processing those exposures.

## Other actions to reduce disk space usage

### Archive Projects

CryoSPARC projects can be archived and moved to a separate device for long-term storage. Archival can be done after using the Cleanup Data tool to reduce the size of the project to its minimum. See:

{% embed url="<https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/guide-data-management-in-cryosparc-v4.0+#archive>" %}

### Reduce database size (v4.3+)

Often after periods of heavy use, the CryoSPARC database (which is separate from project and job directories) can become large. It is possible to reduce the database size with additional steps:

{% content-ref url="/pages/1LzAjuvhV2nWupJX9Gmd" %}
[Guide: Reduce Database Size (v4.3+)](/setup-configuration-and-management/software-system-guides/guide-reduce-database-size-v4.3)
{% endcontent-ref %}

## Data Cleanup Use Cases

The following are example use cases covering recommended actions for clearing project data during various stages of data processing.

### Use Case: Project is not completed, but disk space is running low, or project size must be kept to a minimum

Recommended Actions:

* Restack particles to consolidate useful particle data separately from data that will not be immediately useful in further processing, e.g. junk particles or particles not selected after 2D/3D classification
* Use the project cleaning tool to clear deterministic jobs, including extraction jobs but **not** restack particles jobs.
* Compact all completed CryoSPARC Live sessions (ensuring that useful particles are restacked before compaction)
* Clear unneeded branches of the workspace job tree by marking useful jobs as **final** results and running the Cleanup Data tool

After these actions:

* You can continue processing downstream jobs using the restacked particles
* If you need to re-do upstream processing (e.g. particle picking), pre-processing jobs that were cleared can be re-run to reproduce their outputs

### Use Case: Project is completed and will not be needed in the near future, but there is limited long term storage space

Recommended Actions:

* Annotate unneeded branches of the workspace job tree by marking useful jobs as **final** results
* Run the Cleanup Data tool, and use it to clear all pre-processing jobs (including particle extraction and restack jobs) as well as to clear and delete non-final jobs
* Compact all Live sessions
* Archive the project, allowing results to still be browsed in the CryoSPARC UI but allowing the project directory to be moved
* Move the project directory to long term storage

After these actions:

* If the project ever needs to be restored, move the project directory back into your projects folder and unarchive the project
* Any cleared results can be restored by re-running cleared jobs or restoring Live sessions, albeit at the cost of time


# Guide: Reduce Database Size (v4.3+)

A guide on reducing the size of large CryoSPARC databases using methods provided by MongoDB.

CryoSPARC uses MongoDB to store records and image data. As the size of a CryoSPARC instance grows, the size of the database files can get quite large and affect the performance of the instance. This guide shows how to reduce the size of large databases using methods provided by MongoDB. There are two options.

## Option 1: Compact the database

After clearing or deleting significant amounts of data in CryoSPARC, the MongoDB database can be compacted to potentially reduce the size of the database files. Using the `cryosparcm compact` command will run MongoDB’s `compact` administration command on collections in the database. Be aware, the result of MongoDB’s `compact` varies and it is **not** **guaranteed** to reduce the size of any files in the database.

Please refer to the [MongoDB documentation](https://www.mongodb.com/docs/v3.6/reference/command/compact) for more details on the `compact` procedure.

### Compact database steps

1. Turn on [Maintenance Mode](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/guide-maintenance-mode-and-configurable-user-facing-messages)
2. Allow all previously running jobs on the instance to run to completion, then ensure that there are no running jobs on the instance
3. Shutdown CryoSPARC with `cryosparcm stop`
4. Start the database with `cryosparcm start database`
5. Create a backup of the database with [`cryosparcm backup`](https://guide.cryosparc.com/setup-configuration-and-management/management-and-monitoring/cryosparcm#cryosparcm-backup) . Ensure there is enough space for the backup inside the database directory, or redirect the backup to other storage.
6. Run `cryosparcm compact`

### Verify compaction using a backup (Optional but recommended)

MongoDB’s backup command will write all entries and saved files in the database to a single backup file. The size of this backup file can be used as a reference for how much space the database should take up at minimum. After running `cryosparcm compact`, verify that the size of the database files after compaction is similar to the size of a backup created immediately before or after compaction. If the size of the backup is similar to the size of the database files, it is likely that no further space can be reclaimed from the database. Otherwise, if the size of the backup differs significantly from the size of the database files, `cryosparcm restore` can be attempted to more forcefully reduce database size.

## Option 2: Using backup and restore to reduce database size

If `cryosparcm compact` does not successfully reduce database size, and the size of a database backup is significantly less than the size of the database files after compaction, `cryosparcm restore` can be used to re-write database files and hopefully create a database with a size closer to the backup.

{% hint style="danger" %}
Note: Creating a database backup and restoring from that backup will require free disk space to store both the backup and the restored database. A worst-case estimate of the required free space would be double the size of the current database files. Make sure enough space is available before beginning this process.
{% endhint %}

### Backup and restore steps

1. Turn on [Maintenance Mode](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/guide-maintenance-mode-and-configurable-user-facing-messages)
2. Allow all previously running jobs on the instance to run to completion, then ensure that there are no running jobs on the instance
3. Create a backup of the database with [`cryosparcm backup`](https://guide.cryosparc.com/setup-configuration-and-management/management-and-monitoring/cryosparcm#cryosparcm-backup)
4. Stop CryoSPARC with `cryosparcm stop`
5. Create a new (empty) directory for the restored database, e.g. `mkdir cryosparc_database_from_backup`
6. Change the `CRYOSPARC_DB_PATH` environment variable in `cryosparc_master/config.sh` to be the path to the new (empty) database directory. (This can be changed back to the original database directory if anything goes wrong in the next steps)
7. Run [`cryosparcm restore --file=<path_to_backup_file>`](https://guide.cryosparc.com/setup-configuration-and-management/management-and-monitoring/cryosparcm#cryosparcm-restore). This will restore the backup to the new empty directory that was created and pointed to above.
8. Start CryoSPARC with `cryosparcm start` and verify the restoration was successful by checking that existing projects and jobs were restored correctly.

### After restoring

After a successful database restoration, the old database folder and database backup file can be deleted. The new database directory can also be renamed to its original name, but remember to edit the `CRYOSPARC_DB_PATH` environment variable in `cryosparc_master/config.sh` to match.


# Guide: CryoSPARC Live Session Data Management (≤v4.7)

How to manage the data created by your cryoSPARC Live Sessions via the user interface and data management API.

{% hint style="danger" %}
The Live Session Data Management panel is deprecated and is no longer available in CryoSPARC v5.0+.
{% endhint %}

{% hint style="warning" %}
The images in this section depict CryoSPARC ≤v3.3. For CryoSPARC v4.0+, please see: [Managing Data](/application-guide/managing-data)
{% endhint %}

## Overview

Live Session Data Management tools are available in CryoSPARC v3.0.0+.

To access the Session Management page, click on the "Manage Data" button at the top of the Browse Sessions Page.

![CryoSPARC Live Session Management page.](/files/-MNeyZLU3ZKMpWRAU8qy)

### Metadata Available

You will be greeted with a table that shows an overview of all Projects and Live Sessions including:

* Project and Session UID
* Project or Session title
* Status of the Session (e.g., Running, Paused, Marked Completed)
* File sizes of the following data categories (Projects show the total across all Sessions within that project)
  * Raw import data
  * Motion corrected micrographs
  * Exposure thumbnails
  * Extracted particles
  * Metadata
* Date and time the Project or Session was created
* Date and time the session was last Paused (if applicable)
* Date and time the session was Marked Completed (if applicable)
* Project actions:
  * Refresh statistics for all sessions within a Project (update file sizes)
  * Download a list of all Project statistics (all files linked to each session, organized into groups)
* Session actions:
  * Refresh statistics for a particular Session (update file sizes)
  * Navigate to the Live Session interface
  * *Additional actions than can be performed on a category of data within a session is described below*

{% hint style="warning" %}
Upon first visit to the Data Management page, you will need to click the "Refresh Project Stats" or "Refresh Session Stats" action to populate the file sizes of each category.
{% endhint %}

{% hint style="warning" %}
You can only perform session actions on sessions that are in `Completed` status. To mark a session as `completed`, navigate to the session, and click the `Mark Completed` button in the `Session Information` sub-tab under the `Configuration` tab.
{% endhint %}

## Actions

### Project Actions

You can download a list of file paths (JSON format) of all sessions within a project (organized into categories) by clicking on the 'Download Project Stats' button in the 'Action' column.

### Session Actions

You can perform actions on all five categories in completed sessions by clicking a cell in the table:

![Clicking on a cell will open up the actions menu for each data type.](/files/-MNeydIHiLQ4zuVYlX0x)

1. **Download a list of file paths** (JSON format) for a particular category (such as motion corrected micrographs) via the browser for use in an external archiving utility
   * Available for all categories except for *thumbnails*
2. **Mark a category as 'archived', 'archiving', or 'active'**
   * Available for all categories except for *thumbnails*
3. **Mark a category as 'deleted'** (for when you have used an external tool to delete all files associated with that category)
   * Available for all categories except for *thumbnails*
4. **Delete data from a particular category** (CryoSPARC will delete all files associated with the category)
   * Available for all categories except for *raw data*
   * A confirmation dialog will be presented before any data is deleted:

![](/files/-MNeyhxtQ-30iQsLWDgc)

### Interface Features

Right-clicking over any cell other than file size data will present will three options:

![Interface features available](/files/-MNeymIykQzluflN7EBY)

1. Refresh table data
2. Expand all projects (show all Sessions across all Projects)
3. Collapse all Projects (show only Project totals)

## State Change Hooks

As of v3.3+, CryoSPARC Live Data Management supports executing a script upon a datatype's state change.

To use this feature, add the following environment variables to cryosparc\_master/config.sh:

```bash
export CRYOSPARC_LIVE_DATA_MANAGEMENT_SCRIPT_ENABLE=true
export CRYOSPARC_LIVE_DATA_MANAGEMENT_SCRIPT_PATH=/abs/path/to/script.sh
```

Then, restart CryoSPARC (cryosparcm stop && cryosparcm start).

The trigger will execute the script that you've specified and pass the following arguments: `project_uid`, `session_uid`, `datatype`, `status`

For example, the "Mark as Archiving" button is clicked for micrographs in P45 S1, the script will be triggered with `'P45'`, `'S1'`, `'micrographs'`, `'archiving'`

![](/files/ZlyKi3NxQ27t4pn6gkuw)

Your script can look something like this (the following script reacts to "`archiving`" state changes, does some work, then updates the status of the datatype to "`archived`" once completed):

```bash
#!/bin/sh

# set path to cryosparcm
cryosparcm=/fast5/userhome/sarulthasan/software/cryosparc/cryosparc_master/bin/cryosparcm
# get variables
project_uid=$1
session_uid=$2
datatype=$3
status=$4

echo "Running data management script on "$CRYOSPARC_MASTER_HOSTNAME:$CRYOSPARC_BASE_PORT
echo "cryosparcm: $cryosparcm";
echo ""
echo "project_uid: $project_uid";
echo "session_uid: $session_uid";
echo "datatype: $datatype";
echo "status: $status";
echo ""

datatype_size=$(${cryosparcm} rtpcli "get_datatype_size(project_uid = '$project_uid', session_uid = '$session_uid', datatype = '$datatype')"
datatype_filepaths_json=$(${cryosparcm} rtpcli "get_datatype_file_paths(project_uid = '$project_uid', session_uid = '$session_uid', datatype = '$datatype')")

echo "Total size of $datatype datatype in $project_uid $session_uid is $datatype_size bytes"
echo ""
# echo "All $datatype filepaths: "
# echo $datatype_filepaths_json

if [ $status = 'archiving' ]; then

		# do something with the filepaths here

		# update the status for this data type
    echo "Changing data management state for $datatype in $project_uid $session_uid to 'archived'"
    ${cryosparcm} rtpcli "change_session_data_management_state(project_uid = '$project_uid', session_uid = '$session_uid', datatype = '$datatype', status = 'archived')"
    RESULT=$?
    if [ $RESULT -eq 0 ]; then
        echo "SUCCESS changed data management state for $datatype in $project_uid $session_uid to 'archived'"
    else
        echo "FAILED changing data management state for $datatype in $project_uid $session_uid to 'archived'. Exiting"
        exit 1;
    fi
fi
```

### CryoSPARC Live Data Management API

Here is the rest of the CryoSPARC Live Data Management API that is available via `cryosparcm rtpcli` or `cryosparcm icli`:

* `get_datatype_file_paths(project_uid, session_uid, datatype)`
  * *Get all the file paths associated with a specific datatype inside a session as a json dictionary.*
* `delete_live_datatype(project_uid, session_uid, datatype, filepaths_to_delete=None, user_id=None, asynchronous=True)`
  * *Delete a specific datatype inside a session. Alternatively, provide a json dict of file paths to delete.*
* `get_datatype_size(project_uid, session_uid, datatype)`
  * *Get the total size of a datatype inside a session in bytes.*
* `update_session_datatype_sizes(project_uid, session_uid)`
  * *Updates the session's 'data\_management' top-level key with the current size of each datatype. Additionally returns the entire size of the session's datatypes.*
* `update_all_sessions_datatype_sizes(project_uid)`
  * *Loops through each session in the project and updates all datatype sizes in each session document.*
* `get_data_management_stats(project_uid)` \*\*
  * *Returns a json formatted dictionary that includes the data\_management dictionary of all sessions in the project.*
* `change_session_data_management_state(project_uid, session_uid, datatype, status)`
  * *Modify the data management status for a specific datatype inside a session. Note that this function will call trigger\_live\_data\_management\_script once completed*
* `trigger_live_data_management_script(project_uid, session_uid, datatype, status=None, script_location=None, script_log_path_abs=None)`
  * *Execute the specified script using the parameters sent. Will log all stdout & stderr into a file. Will only execute if environment variable `CRYOSPARC_LIVE_DATA_MANAGEMENT_SCRIPT_ENABLE` is set in cryosparc\_master/config.sh*

## Enabling Access to the Tool

All CryoSPARC users are able to:

1. View the data management table and refresh file sizes via the 'Action' column → 'Refresh Project Stats'/'Refresh Session Stats'
2. Download the JSON list of files for a category within a session or across a project

In order to mark a category as active/archiving/archived/deleted or delete data, users must have access enabled in the admin panel (in the main cryoSPARC web application):

![](/files/-MNeyphZAV1UWiqhl0oH)

Admin users can click on the button within the 'Live Data Management' column to toggle the ability for a user to modify Live session data:

![Change whether a user can execute actions in the session management page.](/files/-MNeysFDyHmApUcjXBCr)


# Guide: Instance Recovery (v5.0+)

How to recover a CryoSPARC instance if the database directory is corrupted or lost.

## The CryoSPARC Database

CryoSPARC uses a database engine ([MongoDB](https://www.mongodb.com/)) under the hood to store information about the CryoSPARC instance, projects, jobs, users, etc. MongoDB stores this data on the filesystem, in a database directory that is created [when CryoSPARC is installed](https://guide.cryosparc.com/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc#install-the-cryosparc_master-package). The contents of the database are necessary for CryoSPARC to function.

<figure><img src="/files/cmaBpI82UEC7DYhkdZk5" alt=""><figcaption></figcaption></figure>

The above diagram shows how data is managed in CryoSPARC during normal operation.

If for some reason, the CryoSPARC database files become lost or corrupt, information from the project directories and instance configuration export can be used to recover the contents of the database, if the following are true:

1. project directories are intact, and
2. a recent instance configuration export file is available (these are produced automatically every 60 minutes by default in CryoSPARC v5.0+ and stored in the `cryosparc_master/run` directory).

This guide describes new utilities in CryoSPARC v5.0+ that make it easy to recover a CryoSPARC instance in this scenario.

{% hint style="info" %}
Consider regular backups of the `cryosparc_master/run/cryosparc_instance_config_*.tar` file. The backup's schedule and location should be chosen such that a recent backup is likely available in case the `cryosparc_master/` directory is disrupted or unavailable.
{% endhint %}

{% hint style="info" %}
If for some reason, project directories are intact but a recent instance configuration export file is **not** available, it can still be possible to recover the instance manually by installing a new CryoSPARC instance, creating users, and then re-attaching all projects that were previously attached to the lost instance.
{% endhint %}

The next section provides the detailed steps to recover from a loss or corruption of the CryoSPARC database directory. The final section of this page gives the steps that can be used to migrate a CryoSPARC instance to new infrastructure, in a scenario where the database directory cannot be preserved.

## Recovery from a lost or corrupt database

<figure><img src="/files/AhxvnhpYsJ1pXTsSQ7iR" alt=""><figcaption></figcaption></figure>

During instance recovery, information from the project directories and the instance configuration export file are combined to recreate the contents of the CryoSPARC database.

### Prerequisites

1. Intact project directories mounted on the CryoSPARC master computer at the same path(s) as before the database loss
2. A recent instance configuration export file from the original instance (e.g. `cryosparc_instance_config_2025_09_29_13h49.tar`)
   * Instance configuration files are exported once per hour to `cryosparc_master/run/` while CryoSPARC is running
   * It is best to use the latest export available to minimize discrepancy between the exported configuration files and the previous database data before it was lost.
3. Reliable storage with enough space for the recovered database, allowing for future growth, mounted on the CryoSPARC master computer.
4. `cryosparc_master/` installation with the same version as the original instance at the time of database loss
5. `cryosparc_worker/` installation(s) with the same version as the `cryosparc_master/` installation

### Steps

{% hint style="warning" %}
This process may take a significant amount of time for instances with many/large projects.
{% endhint %}

1. If CryoSPARC is running, stop CryoSPARC with `cryosparcm stop`.
2. Confirm, using Linux process query and management commands, that CryoSPARC has been [stopped completely](https://guide.cryosparc.com/setup-configuration-and-management/troubleshooting#incomplete-cryosparc-shutdown).
3. Locate the latest available configuration export file from the original instance. In CryoSPARC v5.0+, an updated version of this file is produced every 60 minutes while an instance is running, and is stored in the `cryosparc_master/run` directory. The filename will be similar to `cryosparc_instance_config_2025_09_29_13h49.tar` .
4. **Make a copy of the latest available configuration export file, placing the copy in a location outside the original CryoSPARC installation, for example: `/home/cryosparcuser/latest_cryosparc_instance_config.tar`.** It is important to make a copy so that the export file is retained even if the `cryosparc_master` installation directory is modified or deleted in later steps.
5. Create a new folder for a new recovered database. e.g.

   ```bash
   mkdir /path/to/recovered_cryosparc_database
   ```
6. Change `CRYOSPARC_DB_PATH` in `cryosparc_master/config.sh` to point to the absolute path of the new database folder.

   ```
   export CRYOSPARC_DB_PATH="/path/to/recovered_cryosparc_database"
   ```
7. Start CryoSPARC with `cryosparcm start`. This will result in an empty CryoSPARC instance since the database is empty. You do not need to create users or take any action in the UI at this point.
8. Run the command

   ```bash
   cryosparcm recover -f /home/cryosparcuser/latest_cryosparc_instance_config.tar
   ```

**What does `cryosparcm recover` command do?**

* The recovery process will restore the instance configuration from the exported file, and then attach all projects that were attached at the time the export was created.
* Recovery of some projects may be skipped if an error is encountered while attaching it. Projects that are not recovered can be attached to the instance after the recovery process is complete.
* Project documents marked as deleted will be skipped.
* Project documents marked as detached will be restored to the database so that their records can be viewed in the user interface.
* Project documents marked as archived will be saved as detached instead, as they can no longer be unarchived. Unarchiving is no longer possible as the recovered database does not include processing metadata until those metadata are imported to the database during project attachment. Previously archived projects can be restored to the instance by attaching the original project directory after the recovery process is complete.
* A project recovery report will be generated and can be used to verify the status of projects in the instance after the recovery process completes.

**Example terminal output**

```
user@cryoem:~/cryosparc/cryosparc_package/cryosparc_master$ cryosparcm recover -f instance_config_exports/cryosparc_instance_config_2025_09_24_11h31.tar 
Are you sure you want to recover from 'instance_config_exports/cryosparc_instance_config_2025_09_24_11h31.tar?'
This will replace the current instance configuration and attempt to re-attach projects. [y/N]: y
2025-09-24T11:58:09.157-0400    connected to: cryoem0:62041
2025-09-24T11:58:09.157-0400    dropping: meteor.config
2025-09-24T11:58:09.240-0400    imported 68 documents
Completed mongo import of collection 'config' from instance_config_exports/cryosparc_instance_config_2025_09_24_11h31.tar
2025-09-24T11:58:09.259-0400    connected to: cryoem0:62041
2025-09-24T11:58:09.259-0400    dropping: meteor.sched_config
2025-09-24T11:58:09.267-0400    imported 3 documents
Completed mongo import of collection 'sched_config' from instance_config_exports/cryosparc_instance_config_2025_09_24_11h31.tar
2025-09-24T11:58:09.278-0400    connected to: cryoem0:62041
2025-09-24T11:58:09.279-0400    dropping: meteor.users
2025-09-24T11:58:09.300-0400    imported 13 documents
Completed mongo import of collection 'users' from instance_config_exports/cryosparc_instance_config_2025_09_24_11h31.tar

✓ Completed import of instance configuration

[11:58:10] Skipping inactive project at /storage/CS/projects/recovery_test_projects/CS-recovery-test-project-2 - deleted=False, archived=True,      recover.py:103
           detached=False                                                                                                                                                
           Could not find project document /storage/CS/projects/recovery_test_projects/CS-recovery-test-project-3/project.json                       recover.py:83
           Cannot attach project folder /storage/CS/projects/recovery_test_projects/CS-recovery-test-project-4, the project folder is already       recover.py:116
           locked to another instance with ID 249c59c4-8e41-4019-8415-2bd6bebf4ca1.                                                                                      
Attaching 4 projects... ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 100% 0:00:00
                                                                            Project recovery                                                                             
┏━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┓
┃ UID ┃ Title                   ┃ Project Dir.                                     ┃ Created At                       ┃ Recovery Status                                 ┃
┡━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┩
│ P1  │ Recovery Test Project 1 │ /storage/CS/projects/recovery_test_projec        │ 2025-09-23 20:09:58.250000+00:00 │ Attached                                        │
│     │                         │ ts/CS-recovery-test-project-1                    │                                  │                                                 │
│ P2  │ Recovery Test Project 2 │ /storage/CS/projects/recovery_test_projec        │ 2025-09-24 14:32:20.649000+00:00 │ Saved (Did not attach inactive project)         │
│     │                         │ ts/CS-recovery-test-project-2                    │                                  │                                                 │
│ P3  │ Recovery Test Project 3 │ /storage/CS/projects/recovery_test_projec        │ 2025-09-24 15:06:25.133000+00:00 │ Failed (Could not access project document)      │
│     │                         │ ts/CS-recovery-test-project-3                    │                                  │                                                 │
│ P4  │ Recovery Test Project 4 │ /storage/CS/projects/recovery_test_projec        │ 2025-09-24 15:26:18.011000+00:00 │ Failed (Project directory locked to another     │
│     │                         │ ts/CS-recovery-test-project-4                    │                                  │ instance)                                       │
└─────┴─────────────────────────┴──────────────────────────────────────────────────┴──────────────────────────────────┴─────────────────────────────────────────────────┘
⚠ Writing project recovery results to /u/user/cryosparc/cryosparc_package/cryosparc_master/run/recovery_results_2025_09_24_15h56.json
✓ Project recovery complete - attached 1/4 projects
Synchronizing UID counters...
✓ Instance recovery complete.
```

* Example report

  ```json
  [
      {
          "uid": "P1",
          "title": "Recovery Test Project 1",
          "project_dir": "/storage/CS/projects/recovery_test_projects/CS-recovery-test-project-1",
          "created_at": "2025-09-23 20:09:58.250000+00:00",
          "status": "Attached"
      },
      {
          "uid": "P2",
          "title": "Recovery Test Project 2",
          "project_dir": "/storage/CS/projects/recovery_test_projects/CS-recovery-test-project-2",
          "created_at": "2025-09-24 14:32:20.649000+00:00",
          "status": "Saved (Did not attach inactive project)"
      }
  ]
  ```

After running the above steps successfully, the instance is ready for use.

## Migration of a CryoSPARC instance when the previous database directory is inaccessible

This section describes how to migrate a defunct CryoSPARC instance to a new instance (potentially on a new host) when the previous instance’s database directory has been lost or is inaccessible.

{% hint style="danger" %}
If the previous instance’s database directory **is** available and intact, then it is simpler to copy the previous database directory to the new instance directly, instead of performing instance recovery. See [these instructions for more details](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/tutorial-migrating-your-cryosparc-instance#d.-hosting-cryosparc-on-another-machine) in that case.
{% endhint %}

{% hint style="warning" %}
Regardless of whether a CryoSPARC instance is being migrated using an intact copy of the database or using an instance config export file, one needs to ensure that either

1. the new `cryosparc_worker/` installation paths match records in the database or the config export file or
2. `cryosparc_worker_bin` records are adjusted as needed in the CryoSPARC database after migration and before enqueueing jobs.
   {% endhint %}

### Prerequisites

1. Intact project directories mounted on the CryoSPARC master computer at the same path(s) as before the database loss
2. An instance configuration export archive (e.g.`cryosparc_instance_config_2025_09_29_13h49.tar`)

   copied from the `cryosparc_master/run/` directory of the old CryoSPARC installation after the old installation has been shut down permanently
3. Reliable storage with enough space for the database, allowing for future growth, mounted on the CryoSPARC master computer.
4. \[Optional] Copies of the `config.sh` files from the old `cryosparc_master/` and `cryosparc_worker/` directories. One may refer to these files for custom settings that may continue to apply after migration.

### Steps

1. If needed, install CryoSPARC, but do not start CryoSPARC.

{% hint style="warning" %}
Installation can be skipped if the old installation directories are available on the new CryoSPARC master and worker computers *under the same absolute paths* as prior to the migration.
{% endhint %}

{% hint style="warning" %}
\[If needed] In case the old `cryosparc_worker/` installation is not available on the workers of the new installation, installing worker software under the same path as before removes the need to update scheduler targets’ `worker_bin_path`.
{% endhint %}

2. Follow the [steps](https://www.notion.so/Guide-draft-Instance-recovery-2964324a3a7b80ee9fe8d8b853fca3aa?pvs=21) of the *Recovery from a lost database* section above.
3. \[If needed] In case the scheduler target paths have changed, update the scheduler target configuration with the correct paths.
4. \[If needed] In case you have custom settings in your original instance `config.sh` file that should be kept, copy those custom settings over to the new installation’s `config.sh` .


# Guide: Migrating your CryoSPARC Instance

A guide to moving CryoSPARC from one location to another.

## Introduction

There may come a time when you want to move your CryoSPARC instance from one location to another. This may be between folders, different network storage locations, or even different host machines entirely. There are four main areas we will focus on:

1. The paths of any raw particle, micrograph, or movie data imported into CryoSPARC
2. All CryoSPARC project directories
3. The CryoSPARC database and its (ne&#x77;*)* location
4. The identities/hostnames of compute nodes or the master node and the CryoSPARC binaries

All four of the above areas can be taken care of in isolation, but if **combined**, will amount to a full-out migration of your CryoSPARC instance.

## Requirements

We will be using a combination of the shell as well as an interactive python session to complete this migration. You will need access to the master node in which the CryoSPARC system is hosted.

It is also recommended that a database backup is created before starting anything, and to not use CryoSPARC after the backup is created until the migration is complete. More details are at: [Setup, Configuration and Management](/setup-configuration-and-management/hardware-and-system-requirements)

## Use Cases

### A. Moving only raw particle, micrograph or movie data already imported into CryoSPARC

When raw data is imported into a CryoSPARC project, rather than copy the data into the project directory, symlinks are created inside the import job directories pointing to the original data files. Read [/pages/F3KBgDxkuaoVRFwpV0KW#7.-imported-data-and-symlinks-in-project-directories](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/pages/F3KBgDxkuaoVRFwpV0KW#7.-imported-data-and-symlinks-in-project-directories "mention")for more details.

If you're moving data that you used an "Import Particles", "Import Micrographs" or "Import Movies" job to bring into CryoSPARC, you will need to repair these jobs. When CryoSPARC imports these three types of data, it creates [symlinks](https://devdojo.com/tutorials/what-is-a-symlink) to each file inside the job's `imported` directory. These symlinks may become broken if the original path to the file no longer exists. You can check the status of the symlinks by running `ls -l` inside the `imported` directory of the job. Note: The "Import Templates" and "Import Volumes" jobs copy the specified files directly into the job directory.

#### Modify Project/Job symlinks

{% tabs %}
{% tab title="CryoSPARC v5.0+" %}
Start up an interactive python session

```bash
cryosparcm icli
```

Use the cli to find all the symlinks for an entire project or a single job.

```bash
>>> api.projects.get_symlinks(’P1’)

[{'exists': True,
'link_path': '/bulk8/data/dev_nwong_projects/P1/J2/imported/004525579726026751633_14sep05c_c_00003gr_00014sq_00011hl_00003es.frames.tif',
'link_target': '/bulk8/data/dev_nwong_projects/testdata/empiar_10025_subset/14sep05c_c_00003gr_00014sq_00011hl_00003es.frames.tif'},
…]
```

where

* `link_path` is the path to the symlink file
* `link_target` is the file the symlink points to
* `exists` indicates if the target file exists

Use the command `api.jobs.update_directory_symlinks(project_uid, job_uid, prefix_cut, prefix_new)` where `prefix_cut` is the beginning of the link you'd like to cut (e.g. `/data/EMPIAR`) and where `prefix_new` is what you'd like to replace it with (e.g. `/data`). This function will loop through every file inside the job directory, find all symlinks, and only modify them only if they start with `prefix_cut`. The function returns the number of links it modified. Below it is used in a loop to modify all jobs across all projects all at once.

```python
>>> jobs = api.jobs.find(project_uid="P1")
>>> failed_jobs = []
>>> for job in jobs:
        try:
            print(f"Repairing {job.full_uid}")
            modified_count = api.jobs.update_directory_symlinks(job.project_uid, job.uid, '/data/EMPIAR', '/data')
            print(f"Finished. Modified {modified_count} links.")
        except Exception as e:
            failed_jobs.append((job.project_uid, job.uid))
            print(f"Failed to repair {job.full_uid}: {str(e)}")
...
    
>>> failed_jobs
[]
```

{% endtab %}

{% tab title="CryoSPARC v4.0-v4.7.1" %}
Start up an interactive python session

```bash
cryosparcm icli
```

Use the cli to find all the symlinks for an entire project or a single job.

```bash
>>> cli.get_project_symlinks(’P1’)

or

>>> cli.get_job_symlinks(’P1’, ‘J3’)

[{'exists': True,
'link_path': '/bulk8/data/dev_nwong_projects/P1/J2/imported/004525579726026751633_14sep05c_c_00003gr_00014sq_00011hl_00003es.frames.tif',
'link_target': '/bulk8/data/dev_nwong_projects/testdata/empiar_10025_subset/14sep05c_c_00003gr_00014sq_00011hl_00003es.frames.tif'},
…]
```

where

* `link_path` is the path to the symlink file
* `link_target` is the file the symlink points to
* `exists` indicates if the target file exists

Use the command `cli.job_import_replace_symlinks(project_uid, job_uid, prefix_cut, prefix_new)` where `prefix_cut` is the beginning of the link you'd like to cut (e.g. `/data/EMPIAR`) and where `prefix_new` is what you'd like to replace it with (e.g. `/data`). This function will loop through every file inside the job directory, find all symlinks, and only modify them only if they start with `prefix_cut`. The function returns the number of links it modified. Below it is used in a loop to modify all jobs across all projects all at once.

```python
>>> failed_jobs = []
>>> for job in jobs:
        try:
            print("Repairing %s %s" % (job['project_uid'], job['uid']))
            modified_count = cli.job_import_replace_symlinks(job['project_uid'], job['uid'], '/data/EMPIAR', '/data')
            print("Finished. Modified %d links." % (modified_count))
        except Exception as e:
            failed_jobs.append((job['project_uid'], job['uid']))
            print("Failed to repair %s %s: %s" % (job['project_uid'], job['uid'], str(e)))
...
    
>>> failed_jobs
[]
```

{% endtab %}

{% tab title="CryoSPARC ≤v3.3" %}
Start up an interactive python session

```
cryosparcm icli
```

Execute a MongoDB query for all potentially affected "import" jobs

```python
>>> jobs = list(db.jobs.find({'deleted': False, 'job_type': {'$in': ['import_particles', 'import_movies', 'import_micrographs']}}, {'_id': 0, 'project_uid': 1, 'uid': 1, 'job_type': 1}))
>>> print(jobs)
[{'job_type': 'import_movies', 'project_uid': 'P1', 'uid': 'J3'},
{'job_type': 'import_movies', 'project_uid': 'P2', 'uid': 'J3'},
{'job_type': 'import_movies', 'project_uid': 'P1', 'uid': 'J41'},
{'job_type': 'import_particles', 'project_uid': 'P1', 'uid': 'J42'},
{'job_type': 'import_movies', 'project_uid': 'P2', 'uid': 'J29'},
{'job_type': 'import_particles', 'project_uid': 'P2', 'uid': 'J51'},
{'job_type': 'import_movies', 'project_uid': 'P2', 'uid': 'J55'},
{'job_type': 'import_movies', 'project_uid': 'P2', 'uid': 'J64'},
...]
```

From this point, you can take a look into each list job's `imported` directory

```python
>>> cli.get_job_dir_abs('P1', 'J3')
'/data/cryosparc_projects/P1/J3'
    
>>> !ls -l /data/cryosparc_projects/P1/J3/imported
total 99
lrwxrwxrwx 1 cryosparcuser cryosparcuser 83 Jan 31  2018 14sep05c_00024sq_00003hl_00002es.frames.mrc -> /data/EMPIAR/10025/data/14sep05c_raw_196/14sep05c_00024sq_00003hl_00002es.frames.mrc
lrwxrwxrwx 1 cryosparcuser cryosparcuser 83 Jan 31  2018 14sep05c_00024sq_00003hl_00005es.frames.mrc -> /data/EMPIAR/10025/data/14sep05c_raw_196/14sep05c_00024sq_00003hl_00005es.frames.mrc
lrwxrwxrwx 1 cryosparcuser cryosparcuser 83 Jan 31  2018 14sep05c_00024sq_00004hl_00002es.frames.mrc -> /data/EMPIAR/10025/data/14sep05c_raw_196/14sep05c_00024sq_00004hl_00002es.frames.mrc
lrwxrwxrwx 1 cryosparcuser cryosparcuser 83 Jan 31  2018 14sep05c_00024sq_00006hl_00003es.frames
```

Use the command `cli.job_import_replace_symlinks(project_uid, job_uid, prefix_cut, prefix_new)` where `prefix_cut` is the beginning of the link you'd like to cut (e.g. `/data/EMPIAR`) and where `prefix_new` is what you'd like to replace it with (e.g. `/data`). This function will loop through every file inside the job directory, find all symlinks, and only modify them only if they start with `prefix_cut`. The function returns the number of links it modified. Below it is used in a loop to modify all jobs across all projects all at once.

```python
>>> failed_jobs = []
>>> for job in jobs:
        try:
            print("Repairing %s %s" % (job['project_uid'], job['uid']))
            modified_count = cli.job_import_replace_symlinks(job['project_uid'], job['uid'], '/data/EMPIAR', '/data')
            print("Finished. Modified %d links." % (modified_count))
        except Exception as e:
            failed_jobs.append((job['project_uid'], job['uid']))
            print("Failed to repair %s %s: %s" % (job['project_uid'], job['uid'], str(e)))
...
    
>>> failed_jobs
[]
```

{% endtab %}
{% endtabs %}

### B. Moving Only CryoSPARC Project Directories (and all jobs inside them)

For instructions on moving cryoSPARC project directories in v4.0+, see [Guide: Data Management in CryoSPARC (v4.0+)](/setup-configuration-and-management/software-system-guides/guide-data-management-in-cryosparc-v4.0#use-case-moving-a-project-directory-from-one-storage-location-to-another)

<details>

<summary>Instructions for CryoSPARC ≤ v3.3</summary>

If you're moving the locations of the projects and their jobs, you will need to point CryoSPARC to the new directory where the projects reside. Jobs inside CryoSPARC are referenced by their relative location to their project directory. This allows a user to specify a new location for the project directory only, rather than each job.

**Update a Single Project**

`cryosparcm cli "update_project('PXX', {'project_dir' : '/new/abs/path/PXX'})"`

* Where `'PXX'` is the project UID and `'/new/abs/path/PXX'` is the new directory.

**Updating Multiple Projects**

**Step One - Identify All Project Directories**

Start up an interactive python session

```
cryosparcm icli
```

Execute a MongoDB query to list all project directories

```python
>>> projects = list(db['projects'].find({}, {'uid': 1, 'project_dir': 1, '_id': 0}))
>>> projects
    [{'project_dir': '/data/cryosparc_projects/P1', 'uid': 'P1'},
     {'project_dir': '/data/cryosparc_projects/P2', 'uid': 'P2'},
     {'project_dir': '/data/cryosparc_projects/P3', 'uid': 'P3'},
     {'project_dir': '/data/cryosparc_projects/P4', 'uid': 'P4'},
    ...]
```

**Step Two - Modify One or Many Project Directories**

Use the command `update_project(project_uid, attrs, operation='$set')` where `attrs` is a dictionary whose keys correspond to the fields in the project document to update. In the following example, `update_project` used in a loop to modify all project directory paths.

```python
>>> failed_projects = []
>>> new_parent_dir = '/cryoem/cryosparc_projects'
>>> for project in projects:
        new_project_dir = os.path.join(new_parent_dir, os.path.basename(project['project_dir']))
        try:
            print("Modifying project directory for %s: %s --> %s" % (project['uid'], project['project_dir'], new_project_dir))
            cli.update_project(project['uid'], {'project_dir': new_project_dir})
        except Exception as e:
            failed_projects.append(project['uid'])
            print("Failed to update %s: %s" % (project['uid'], str(e)))
...

>>> failed_projects
[]
```

</details>

### C. Moving the CryoSPARC database

The CryoSPARC database doesn't necessarily have to be in the same location as the CryoSPARC installation directories. To move the database location, use the following steps.

{% hint style="warning" %}
Ensure that neither CryoSPARC project directories nor the CryoSPARC database are modified while the database is being moved or copied. If the database and project directories become out-of-sync, CryoSPARC may not function correctly and/or project directories may get corrupted.
{% endhint %}

#### Step One - Shut down CryoSPARC

```
cryosparcm stop
```

#### Step Two - Move the Database

```
rsync -r --links /data/cryosparc/cryosparc_database/* /new/path/cryosparc_database
```

#### Step Three - Modify Configurations

Navigate to the `cryosparc_master` directory

```
cd /data/cryosparc/cryosparc_master
```

Modify `config.sh` to contain the new directory path to the database

```
nano config.sh
    
#modify the line below
export CRYOSPARC_DB_PATH="/new/path/cryosparc_database"
```

#### Step Four - Start CryoSPARC again

```
cryosparcm start
```

### D. Hosting CryoSPARC on another machine

If you have a CryoSPARC instance where the master application was running on one machine and now need to run the master on a new machine, use the following steps.

If the CryoSPARC master installation directory (i.e. `cryosparc_master` ) resides on a filesystem that is shared and mounted at the same location on both the old and new machine, use **Option 1** which is simplest. This may be the case for example if CryoSPARC was installed in a user home directory such as `/home/cryosparcuser/cryosparc_master` and home directories are shared across machines in your setup.

If the CryoSPARC installation directory is not on shared storage, use **Option 2**.

{% hint style="warning" %}
In either case, the new machine must have project directories and raw data directories mounted at the same locations as the old machine, and must have access to the same worker nodes and cluster schedulers as the old machine.
{% endhint %}

### Option 1: `cryosparc_master` is on a shared Filesystem

1. Shut down CryoSPARC on the old machine.

```
cryosparcm stop
```

2. On the new machine, **log in as the same user as on the old machine**, and navigate to the CryoSPARC installation directory and modify the configuration:

Navigate to the cryosparc\_master directory

```
cd /data/cryosparc/cryosparc_master
```

Modify config.sh to list the new master node hostname

```
nano config.sh
    
#modify the line below
export CRYOSPARC_MASTER_HOSTNAME="newnode"
```

3. On the new machine, start CryoSPARC using `cryosparcm start`

### Option 2: Installation is not shared

In this case, follow the [Installation guide](/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc) to install a fresh copy of CryoSPARC on the new machine, **using the same LICENSE\_ID as was used on the old machine.** Then:

1. On the old machine, shut down CryoSPARC using `cryosparcm stop` .
2. Copy the CryoSPARC database directory from the old machine to the new machine, for example at `/new/path/cryosparc_database`.
3. On the new machine, shut down CryoSPARC using `cryosparcm stop`.
4. On the new machine, edit the `cryosparc_master/config.sh` file to point to the new database path by editing `export CRYOSPARC_DB_PATH=/new/path/cryosparc_database` .
5. On the new machine, in the `cryosparc_master/config.sh` file, add any configuration variable that were present in the same file on the old machine, if relevant.
6. On the new machine, start CryoSPARC using `cryosparcm start` .


# Deploying CryoSPARC on AWS

Version 1.0 (May 10, 2021)

{% hint style="warning" %}
This deployment guide is based on version 2 of AWS ParallelCIuster. For an updated configuration using a newer version of AWS ParallelCluster, please see [this](https://github.com/aws-samples/cryoem-on-aws-parallel-cluster) AWS sample.
{% endhint %}

{% hint style="info" %}
**This Deployment Guide provides end-to-end sample instructions for deploying** [**CryoSPARC™**](https://cryosparc.com/)**, a state-of-the-art scientific software platform for cryo-EM, on AWS using AWS ParallelCluster.** CryoSPARC is developed by [Structura Biotechnology](https://structura.bio/) Inc. Additional information about CryoSPARC, including [licensing](https://guide.cryosparc.com/licensing), is available at [guide.cryosparc.com](https://guide.cryosparc.com/).
{% endhint %}

## 1. **Introduction**

Cryo-electron microscopy (cryo-EM) is a biophysical technique that allows scientists to determine the structure of biological macromolecules and assemblies. This [technology that was awarded the 2017 Nobel Prize in Chemistry](https://www.lsi.umich.edu/news/2017-10/chilled-proteins-3-d-images-cryo-electron-microscopy-technology-just-won-nobel-prize) uses advanced microscopes to reveal 3D structures of biomolecules in near-native states. Cryo-EM is rapidly becoming the go-to technique for protein structure determination in life-sciences and drug discovery. Just recently, cryo-EM was used to produce [the first atomic-level 3D structure of the spike protein responsible for the COVID-19 virus](https://cryosparc.com/blog/2019-nCoV).

Storing the micrographs (images produced by the microscope) requires enormous data storage and the processing workflow requires massive computing power. This workload is therefore an ideal use case for High-Performance Computing in Amazon Web Services.

![](/files/-M_LaZkFcMdSt_A0Q1pu)

A typical cryo-EM workflow involves biological sample preparation, data collection, and finally computation. Images of a sample are collected with a transmission electron microscope. The raw data, which is generally 5-10 TB in size, is then processed to reconstruct a three-dimensional structure of the protein of interest from the two-dimensional images. Resolving a 3D structure often requires multiple iterations through the entire processing pipeline, or parts of the pipeline, for a near-atomic resolution result. In many cases, the entire workflow starting from data collection must be repeated to achieve state-of-the-art results.

![](/files/-M_LaZkGeOXadzK490ze)

### Benchmarks and Cost Estimates

This guide is accompanied by a Performance Benchmarks document outlining various steps in running a CryoSPARC workflow, timings and cost estimates:

{% content-ref url="/pages/-M\_Lb0wcBRS1s2uu7ThC" %}
[Performance Benchmarks](/setup-configuration-and-management/cryosparc-on-aws/performance-benchmarks-aws)
{% endcontent-ref %}

The benchmark was performed on the [EMPIAR-10288](https://www.ebi.ac.uk/pdbe/emdb/empiar/entry/10288/) dataset of approximately 2800 micrographs totalling 476GB. For the deployment guide, we recommend running the smaller [T20S tutorial](https://guide.cryosparc.com/processing-data/cryo-em-data-processing-in-cryosparc-introductory-tutorial) that ships with CryoSPARC. Approximate costs for the T20S example on different EC2 instance types are:

* p4d.24xlarge: **$37 USD**
* p3dn.24xlarge: **$42 USD**
* g4dn.metal: **$10 USD**

{% hint style="warning" %}
**NOTE: This guide serves as an example of possible installation options, performance and cost, but each user’s results may vary.** Performance and costs may scale differently depending on the specific compute setup, data being processed, how long AWS compute resources are being used, specific steps used in processing, etc.
{% endhint %}

This deployment guide uses [Amazon EC2](https://aws.amazon.com/ec2/), [Amazon Simple Storage Service (S3)](https://aws.amazon.com/s3/), [Amazon FSx for Lustre](https://aws.amazon.com/fsx/), and [Amazon CloudFormation](https://aws.amazon.com/cloudformation/). See the architecture diagram in section 11.3 for details.

## **2. Pre-requisites**

This deployment guide assumes minimal AWS knowledge; however, there are a few prerequisites. The first step is to create an AWS account, as described here:\
<https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/get-set-up-for-amazon-ec2.html#sign-up-for-aws>

You will also need the following:

* A [CryoSPARC license ID](https://guide.cryosparc.com/setup-configuration-and-management/how-to-download-install-and-configure/obtaining-a-license-id)
* A computer with internet running macOS or Linux. For Windows users, a terminal emulator.
* An Internet browser such as Chrome or Firefox
* Familiarity with Linux terminal commands
* Time to Request EC2 Service Quota increases (at least 24 hours)
* IAM Permissions

## **3. AWS Management Console**

To log into the AWS Management Console, navigate to this address: <https://aws.amazon.com/console/> and click on the link labelled “Already have an account? Sign in.” You will be prompted for an Account ID or alias, IAM user name, and Password. If you have just created a new account you will need to sign up with your root user email and password.

Once you are logged into the AWS Management Console, spend some time becoming familiar with the interface. This page is a central place you can use to find and learn about AWS services as well as to manage and monitor your account. Some key items to note are the Search Bar and the AWS region menu. The latter allows you to select the geographical region you would like your AWS resources to be located. AWS currently has 24 regions, each of which contains multiple Availability Zones (AZ), each containing one or more data centers, spread out across the globe where you can run your cryo-EM analysis.

{% hint style="warning" %}
Note that some newer compute instances may not be available in all AZs. Please select the AZ that has all the resources required for your workflow.
{% endhint %}

## **4. AWS CLI**

The [AWS Command Line Interface](https://docs.aws.amazon.com/cli/latest/userguide/cli-chap-install.html) (AWS CLI) is an open-source tool that enables you to interact with AWS services from your command-line shell. Download AWS CLI using the commands below:

### For Linux

```bash
$ curl "https://awscli.amazonaws.com/awscli-exe-linux-x86_64.zip" -o "awscliv2.zip"
$ unzip awscliv2.zip
$ sudo ./aws/install
```

### For macOS

```bash
$ curl "https://awscli.amazonaws.com/AWSCLIV2.pkg" -o "AWSCLIV2.pkg"
$ sudo installer -pkg AWSCLIV2.pkg -target /
```

If you do not have superuser (sudo) permissions, AWS CLI can also be installed using pip with regular user permissions. You can find more instructions [here](https://docs.aws.amazon.com/cli/latest/userguide/install-macos.html#awscli-install-osx-pip).

Verify that the AWS shell was installed correctly with the following commands (example outputs included)

```
$ which aws
/usr/local/bin/aws

$ aws --version
aws-cli/2.0.47 Python/3.7.4 Darwin/18.7.0 botocore/2.0.0
```

### Configure the AWS CLI tool

In order to access services associated with your AWS account, you must first configure the AWS CLI. You will need to provide an Access Key ID and a Secret access key. [For more information on configuring the AWS CLI, see this article.](https://docs.aws.amazon.com/cli/latest/userguide/cli-configure-quickstart.html)

## 5. IAM Role & Permissions Required

Whether you are using a new or existing AWS account to deploy CryoSPARC in the cloud, it’s best to create a new IAM user specifically for this purpose. Doing so allows you to give the IAM user-specific policies that are scoped to the resources and actions needed to complete the deployment.

[This article explains how to create a new IAM user in your AWS account](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_users_create.html#id_users_create_console). The type of access required is “Programmatic Access”, as this IAM user will only be used for the `deploy-cryosparc.sh` script via the CLI.

### Permissions Required

During testing, the following AWS managed policies were attached to the IAM user deploying CryoSPARC:

* AmazonEC2FullAccess
* AmazonFSxFullAccess
* AmazonS3FullAccess
* AmazonDynamoDBFullAccess
* CloudWatchLogsFullAccess
* AmazonRoute53FullAccess
* AWSCloudFormationFullAccess
* AWSLambda\_FullAccess

as well as the following custom managed policy:

```
{
  "Version": "2012-10-17",
  "Statement": [
   {
    "Sid": "CustomIAMCryoSPARCPolicy",
    "Effect": "Allow",
    "Action": [
     "iam:CreateInstanceProfile",
     "iam:DeleteInstanceProfile",
     "iam:GetRole",
     "iam:RemoveRoleFromInstanceProfile",
     "iam:CreateRole",
     "iam:DeleteRole",
     "iam:AttachRolePolicy",
     "iam:PutRolePolicy",
     "iam:AddRoleToInstanceProfile",
     "iam:PassRole",
     "iam:DetachRolePolicy",
     "iam:DeleteRolePolicy",
     "iam:GetRolePolicy"
    ],
   "Resource": "*"
  }
 ]
}
```

[See this article for more details on creating custom IAM policies.](https://docs.aws.amazon.com/IAM/latest/UserGuide/access_policies_create-console.html)

## **6. EC2 Dashboard**

[Amazon EC2](https://aws.amazon.com/free/?all-free-tier.sort-by=item.additionalFields.SortRank\&all-free-tier.sort-order=asc\&awsf.Free%20Tier%20Categories=categories%23compute\&trk=ps_a134p000004f2ZFAAY\&trkCampaign=acq_paid_search_brand\&sc_channel=PS\&sc_campaign=acquisition_US\&sc_publisher=Google\&sc_category=Cloud%20Computing\&sc_country=US\&sc_geo=NAMER\&sc_outcome=acq\&sc_detail=amazon%20ec2\&sc_content=EC2_e\&sc_matchtype=e\&sc_segment=467723097970\&sc_medium=ACQ-P%7CPS-GO%7CBrand%7CDesktop%7CSU%7CCloud%20Computing%7CEC2%7CUS%7CEN%7CText\&s_kwcid=AL!4422!3!467723097970!e!!g!!amazon%20ec2\&ef_id=EAIaIQobChMI0aCFud2A7wIVLh6tBh3ApQB9EAAYASAAEgI3uvD_BwE:G:s\&s_kwcid=AL!4422!3!467723097970!e!!g!!amazon%20ec2) is a service that provides scalable and flexible computing capacity in the AWS cloud. Using Amazon EC2 eliminates the need to invest in computing hardware and enables quick development and deployment of applications. It allows many virtual servers to be configured with security, networking and storage management. EC2 virtual servers, also known as instances, are the building blocks for supercomputing on AWS.

The EC2 dashboard displays resources and provides the ability to launch an instance. On the left-hand side of the dashboard, there are links to EC2 limits, instances, AMIs, Security Groups and keys.

![](/files/-M_LaZkH-uUuXQw-TSXI)

### **6.1. EC2 Service Quotas**

All accounts initially have a lower limit to protect against fraud. Increasing these limits requires a simple request based on your region and what instances you need. Please verify that GPU-based p\* and g\* instances are available in your region. This deployment guide uses the us-east-1 (US East - N. Virginia) region.

From the AWS console, search for Service Quotas (AWS Services – Service Quotas) to go to the Elastic Compute Cloud section. Type in ‘on-demand’ to filter. Select the type of instance for which you wish to increase the vCPU limits and click ‘Request Limit Increase’. For this workload, request a limit increase for p\*, g\*, and c\* instance types. Select the desired region. Enter a case description on why you want the limits to be increased and click Submit. You should receive a response from AWS Support within 24 hours confirming your limits have been increased. You can read more about service limits [here](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-resource-limits.html).

![](/files/-M_LaZkI0pcxMjTIFFsj)

### **6.2. EC2 key**

A key pair, consisting of a private key and a public key, is a set of security credentials required to prove your identity when connecting to an instance. Create an [EC2 key](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-key-pairs.html) pair in the region you plan to deploy the CryoSPARC cluster. This can be done in two ways:

1. Go to the EC2 dashboard in your AWS console. Select Key pairs, and select Create key pair. Give an appropriate name, download the key pair.
2. Open your terminal and type the following command:

```
$ aws ec2 create-key-pair \
--key-name cryoSPARC \
--region name-of-region \
--output text > ~/.ssh/key-cryoSPARC
```

* Substitute name-of-region for the region where you want to deploy your cluster. For this guide, enter `us-east-1`

```
$ chmod 600 ~/.ssh/key-cryoSPARC
```

Your key is stored in the `~/.ssh` directory and is called `key-cryoSPARC`.

## **7. Amazon VPCs and Subnets**

[Amazon Virtual Private Cloud](https://docs.aws.amazon.com/vpc/latest/userguide/what-is-amazon-vpc.html)[ ](https://docs.aws.amazon.com/vpc/latest/userguide/what-is-amazon-vpc.html)(Amazon VPC) can launch AWS resources in a virtual network that you have defined. A VPC is a virtual network dedicated to your AWS account. It is logically isolated from other virtual networks. A subnet is a range of IP addresses in your VPC in which you can launch AWS resources. This virtual network closely resembles a traditional network operating in an on-premises data center, with the benefits of using the scalable infrastructure of AWS.

A network access control list (ACL) is an optional layer of security for your VPC. that acts as a firewall, controlling traffic in and out of one or more subnets. A security group acts as a virtual firewall for your instance to control inbound and outbound traffic. Security groups operate at the instance level, not the subnet level. Therefore, each instance in a subnet in your VPC can be assigned to a different set of security groups.

## 8. Deployment Scripts

{% hint style="success" %}
Before proceeding further, download the zip folder containing all the files required for this deployment here: <https://github.com/cryoem-uoft/aws-deployment-guide/archive/refs/tags/v1.0.zip>
{% endhint %}

The folder contains 5 files:

1. `README.md`
2. `cryosparc-pcluster.config.template`
3. `deploy-cryosparc_v1.sh`
4. `install-cryosparc_v1.sh`
5. `vpc-cryosparc.template`

Optionally, browse the code on GitHub <https://github.com/cryoem-uoft/aws-deployment-guide>**.**

## **9. Amazon S3**

[Amazon Simple Storage Service](https://docs.aws.amazon.com/AmazonS3/latest/gsg/GetStartedWithS3.html) (Amazon S3) is a global object storage service for storing data and securing it from unauthorized access. Use Amazon S3 to store raw images for cryoSPARC to analyze.

A single bucket is required for this guide. In the following instructions, replace the given example bucket name with your own, as bucket names must be globally unique. The name must not contain any uppercase characters. Since the bucket will store raw data, make sure your bucket is private.

{% hint style="warning" %}
S3 buckets and compute resources required to run a cryoSPARC workflow must be in the same AZ (see Section 10.2). Note that some newer compute instances may not be available in all AZs. Please create the bucket in an AZ that has all the resources required for your workflow.
{% endhint %}

Create the bucket **`cryosparc-test-data-np`** for raw data. This S3 bucket will be linked to the Amazon FSx for Lustre service for high-performance read and write storage operations.

To create the bucket and upload the raw `.tif` and `.mrc` movies (multi-frame micrographs) to **`cryosparc-test-data-np`:**

```
$ aws s3 mb s3://cryosparc-test-data-np
$ aws s3 cp ./ s3://crysoparc-test-data-np --exclude "*" --include "*.tif"
$ aws s3 cp ./ s3://cryosparc-test-data-np --exclude "*" --include "*.mrc"
```

From the S3 management console, verify that all files were successfully uploaded to the bucket.<br>

## **10. Amazon FSx for Lustre**

[Amazon FSx for Lustre](https://aws.amazon.com/fsx/lustre/?nc=sn\&loc=0) is a fully managed service that provides cost-effective, high-performance storage for compute workloads. Many workloads such as machine learning, high-performance computing (HPC), video rendering and financial simulations depend on compute instances accessing the same set of data through high-performance shared storage. Amazon FSx for Lustre file systems can also be linked to Amazon S3 buckets, enabling access and process data concurrently from a high-performance file system. Amazon FSx for Lustre can also be configured to back up data to Amazon S3, and further to Amazon S3 Glacier to optimize costs for data backup.

You can choose between two file systems when using Amazon FSx for Lustre:

1. Persistent File Systems - these are ideal for long-term storage.
2. Scratch File Systems - these are ideal for temporary and short-term storage.

This guide uses the Scratch file system. This provides the most performant and cost-effective option. If you require data resilience, consider the persistent deployment guide.

## **11. AWS CloudFormation**

[AWS CloudFormation](https://aws.amazon.com/cloudformation/) provides a way to model a collection of related AWS and third-party resources, provision them quickly and consistently, and manage them throughout their lifecycles, by treating infrastructure as code. A CloudFormation template describes your desired resources and their dependencies so you can launch and configure them together as a stack. You can use a template to create, update and delete an entire stack as a single unit, as often as you need to, instead of managing resources individually. You can manage and provision stacks across multiple AWS accounts and AWS Regions.

The script *`vpc-cryosparc.template`* is a CloudFormation template that deploys the VPC and subnets. As is, the script creates a VPC and two subnets. The subnet for the head instance is public, and the compute instances are placed in a private subnet. Outside the VPC, you can only log into the head instance in the public subnet (via SSH).

## **12. AWS ParallelCluster**

[AWS ParallelCluster ](https://aws.amazon.com/hpc/parallelcluster/)enables you to quickly build an HPC environment on AWS. It automatically sets up the required compute resources, a shared filesystem, and offers a variety of batch schedulers. You define all the resources you need in a config file.

This deployment will use ParallelCluster 2.10.1. CryoSPARC requires the multiple queues feature in ParallelCluster, which is only supported by versions 2.9.0 and later.

To install ParallelCluster 2.10.1, run the following command on your local machine:

```
$ pip3 install aws-parallelcluster --user
```

This enables access to the terminal command-line tool `pcluster`. First, confirm that you have the correct version installed.

```
$ pcluster version
2.10.1
```

### **12.1. Configure the cluster**

Open *`cryosparc-pcluster.config.template`* in a text editor; here, provide the details of the cluster required to deploy and run the CryoSPARC workflow.

An AWS ParallelCluster configuration is defined in multiple sections. A section starts with the section name in square brackets, followed by parameters and configuration. [See this page](https://docs.aws.amazon.com/parallelcluster/latest/ug/configuration.html) for more information about the config file. At the time of deployment, a *`cryosparc.config`* file is created from the *`cryosparc.config.template`* file (more on this later).

Look through the values that are not explicitly defined in *`cryosparc-pcluster.config.template`*.

* `aws_region_name` - the name of the region
* `key_name` - key pair name
* `post_install` - the path to the install script that runs after the cluster is created
* `s3_read_resource` - the S3 bucket with raw data
* `vpc_id` - the VPC id
* `master_subnet_id` - the subnet id where the head node resides
* `compute_subnet_id` - the subnet id where the compute node resides

You will provide the region, key pair name S3 bucket names and CryoSPARC license ID when you deploy the cluster. The *`deploy-cryosparc.sh`* fetches the remaining values about the networking infrastructure created by *`vpc-cryosparc.template`*.

This config file uses an Amazon EC2 c5n.9xlarge instance as the head node. The c5n.9xlarge instance uses an Intel Xeon Platinum (Skylake) processor with 36 vCPUs. All instances use Amazon Linux 2, a Linux server operating system used by Amazon Web Services (AWS), as their operating system. None of the compute instances are running yet; they will run as needed. Three computing queues are also defined: **gpu-large**, **gpu-med**, and **gpu-small**. Each queue can host multiple EC2 instances. For example, the **gpu-large** queue is made of p4d.24xlarge, p3.16xlarge, and g4dn.metal instances. Multiple queues are very useful since different steps of the CryoSPARC workflow run better with different instances. See the attached benchmarking guide for instance recommendations.

### **12.2. How to deploy**

In your local directory, verify that the three provided scripts exist:

* *`cryosparc-pcluster.config.template`*
* *`deploy-cryosparc.sh`*
* *`vpc-cryosparc.template`*

Before launching the cluster, choose the Availability Zone (AZ) where the cluster will run. All EC2 instance types you choose to instantiate should be available within the same AZ in the region your S3 bucket is in. If a required EC2 instance type is not available in any of the AZs in your region, re-create your S3 bucket in a region that has them available.

Run the command below to see where a given instance type is available.

```
$ aws ec2 describe-instance-type-offerings \
--location-type availability-zone \
--region your-aws-region \
--filters Name=instance-type,Values=type of your instance \
--query "InstanceTypeOfferings[*].Location" \
--output text
```

* `--region` Region you want to deploy the cluster in
* `Values=` type of instance

For example - this command outputs the AZs in the us-east-1 region where g4dn.12xlarge instances are available.

```
$ aws ec2 describe-instance-type-offerings \
--location-type availability-zone \
--region us-east-1 \
--filters Name=instance-type,Values=g4dn.12xlarge \
--query "InstanceTypeOfferings[*].Location" \
--output text
us-east-1b us-east-1f us-east-1d us-east-1c us-east-1a
```

Finally, to deploy the cluster

```
$ ./deploy-cryosparc_v1.sh --region your-aws-region \
--cluster-name cryosparc \
--az <your-availability-zone> \
--config-bucket cryosparc-demo-np \
--data-bucket cryosparc-test-data-np \
--aws-key key-cryoSPARC \
--cryosparc-license-id <your-cryosparc-license>
```

* `--region` Region in which to deploy the cluster
* `--cluster-name` Name of the cluster for identification purposes
* `--az` AZ in which to deploy the cluster. Make sure that instance is available in the specified AZ.
* `--data-bucket` Existing S3 bucket created earlier. This will be linked to the EC2 instance with Amazon FSx for Lustre. All movies will be uploaded here.
* `--aws-key` Name (not the path) of the SSH key you created earlier
* `--cryosparc-license-id` The license ID provided by Structura Bio for your cryoSPARC instance

During deployment, a *`cryosparc.config`* file is created from the *`cryosparc.config.template`* file with all the values for the specified variables. This *`cryosparc.config`* is specific to the cluster you just deployed. Check the *`cryosparc.config`* file and ensure all the information is correct. Pay particular attention to the cluster `cryosparc`, `vpc cryosparc-vpc` and `fsx cryosparc-fsx` sections.

Retain the *`vpc-cryosparc.template`* for later use. The deployment will take about 30 minutes.

### **12.3. Deployed Architecture**

The script deploys a VPC in the region you selected with two subnets. A cluster called cryosparc with c5n.9xlarge as the head node is also deployed. The head node resides in the public subnet and the compute node (launched as needed) resides in the private subnet. The head node hosts both the cryoSPARC web interface and the database. Additionally, the script creates and mounts an EBS volume of type gp2 as a shared file system and an Amazon FSx for Lustre file system.

![](/files/-M_Lb3viHHPUbJy91Epp)

### **12.4. Launch CryoSPARC web interface**

As the process completes, navigate to the AWS Management Console and take a look at which resources were deployed. In your terminal, look for instructions to connect to CryoSPARC’s web interface. Following the prompts, log into the head node of the cluster over SSH. The prompt will look like:

![](/files/-M_LaZkK3TH3NzrnOoQ1)

```bash
$ ssh -i /path/to/key/key-cryoSPARC ec2-user@publicIPofyourinstance
```

Once logged in, create a new CryoSPARC user.

```bash
$ cryosparcm createuser \
--email "<youremail@gmail.com>" \
--password "<yourpassword>" \
--username "<yourusername>" \
--firstname "yourname>" \
--lastname "<yourlastname>"
```

The password should not contain special characters. When finished, log out of the head node. Set up an SSH tunnel to the CryoSPARC head node to connect to the CryoSPARC web interface:

```bash
$ ssh -i /path/to/key/key-cryoSPARC -N -f -L \ localhost:45000:localhost:45000 ec2user@publicIPofyourinstance
```

Open a web browser to [http://localhost:45000](http://localhost:45000/) and log into CryoSPARC using the username and password created.

Before you start your production workload, familiarize yourself with the CryoSPARC environment. Follow the instructions [here](https://guide.cryosparc.com/processing-data/cryo-em-data-processing-in-cryosparc-introductory-tutorial) and run your first CryoSPARC workload. More information about CryoSPARC configuration and management is available [here](https://guide.cryosparc.com/setup-configuration-and-management/how-to-download-install-and-configure/cryosparc-cluster-integration-script-examples).

### 12.5. Export output data to Amazon S3

After the workload is completed, download the data from the Amazon FSx for Lustre file system to Amazon S3. This uses a [data repository task](https://docs.aws.amazon.com/fsx/latest/LustreGuide/export-data-repo-task.html).

Use the `create-data-repository-task` command to export data from your Amazon FSx for Lustre file system to back to the data bucket:

```
aws fsx create-data-repository-task \
--file-system-id fs-xxxxxxx \
--type EXPORT_TO_REPOSITORY \
--paths path1 \
--report Enabled=true,Scope=FAILED_FILES_ONLY,Format=REPORT_CSV_20191124,Path=s3://crysoparc-test-data-np/data-repo-report
```

* `--file-system-id` The id for the Lustre file system created. Find this in the AWS console, under Amazon FSx.
* `--paths` The paths of the directory or file you want to export relative to the mount point of the file system. If the mount point is `/mnt/fsx` and `/mnt/fsx/path1` is a directory or file on the file system you want to export, then provide `path1`.

## 13. **Tearing the cluster down**

Unless you plan to run more analysis immediately, we recommend tearing down the cluster to avoid incurring costs. The cluster scales down based on your config file. The current config file ensures the head node and the Amazon FSx for Lustre file system continue running. You can also completely destroy the cluster and spin up a new cluster as needed.

From the AWS command line, execute the following:

```
$ aws --region your-aws-region cloudformation delete-stack --stack-name parallelcluster-cryosparc 
```

to delete the cluster. This includes instances, attached volumes, FSx file system, etc.

```
$ aws --region your-aws-region cloudformation delete-stack --stack-name cryosparc-VPC
```

to delete all networking-related infrastructure. This deletes the VPC and subnets.

These two commands should delete the infrastructure initialized in this guide. However, they retain the S3 buckets with data and installation scripts. Delete the buckets if no longer needed, or archive them for a minimal [price](https://aws.amazon.com/s3/pricing/).

After 15 minutes, log into your AWS console. Check for resources that are no longer required and delete them.

## Release Notes

Version 1.0 (May 10, 2021)


# Performance Benchmarks

Version 1.0 (May 10, 2021)

{% hint style="info" %}
**This Benchmark Guide accompanies the** [**Deploying CryoSPARC on AWS Guide**](/setup-configuration-and-management/cryosparc-on-aws)**. This guide provides an overview of benchmarks performed on a sample cryo-EM data processing workflow in** [**CryoSPARC**](https://cryosparc.com/) **on AWS ParallelCluster. A typical workflow in CryoSPARC involves multiple steps, and it’s important to understand the computational requirements of each step to build a cluster on AWS that is both performant and economical.**
{% endhint %}

In addition to benchmarking results, this guide also presents best practices around EC2 instance selection, CryoSPARC configuration, and file system considerations.

{% hint style="warning" %}
**NOTE: This guide serves as an example of possible installation options, performance and cost, but each user’s results may vary.** Performance and costs may scale differently depending on the specific compute setup, data being processed, how long AWS compute resources are being used, specific steps used in processing, etc.
{% endhint %}

## **Benchmark Data**

The dataset used for benchmarking is [EMPIAR-10288](https://www.ebi.ac.uk/pdbe/emdb/empiar/entry/10288/) (cannabinoid receptor 1-G protein complex). It’s composed of 2756 TIFF images constituting 476 GB of data. The dataset is moderate in size (compared to production cryo-EM workloads) but was chosen because the size meant a large number of simulations could be run to test a range of different architecture options. Later benchmarking efforts will build on the analysis presented here and will be applied to larger datasets.

{% hint style="danger" %}
There are a few files in the raw dataset that need to be discarded before processing, due to a mismatch in shapes when compared to the rest of the dataset. These files are:

* CB1\_\_00004\_Feb18\_23.33.18.tif
* CB1\_\_00005\_Feb18\_23.34.19.tif
* CB1\_\_00724\_Feb19\_12.00.25.tif
  {% endhint %}

{% hint style="warning" %}
The raw dataset comes with two gain reference files in the `.dm4` format, which at this time, CryoSPARC does not support. Use the following file as a gain reference for the entire dataset, which has been converted to `.mrc` from one of the `.dm4` files:

<https://structura-assets.s3.amazonaws.com/files/CountRef_CB1__00826_Feb19_14.19.25.mrc>
{% endhint %}

### Processing Steps

**1. Import Movies**

* Movies data Path: `.../10288/data/*.tif`
* Gain reference path: `...10288/data/CountRef_CB1__00826_Feb19_14.19.25.mrc`
* Flip gain ref in Y? `True`
* Raw pixel size (Å): `0.86`
* Accelerating Voltage (kV): `300`
* Spherical Aberration (mm): `2.7`
* Total exposure dose (e/Å^2): `58`

**2. Patch Motion Correction**

**3. Patch CTF Estimation**

**4. Blob Picker**

* Minimum particle diameter (Å): `100`
* Maximum particle diameter (Å): `150`
* Use circular blob: `True`
* Use elliptical blob: `True`
* Number of micrographs to process: `100`

**5. Extract from Micrographs**

* Extraction Box Size: `320`
* Fourier crop to box size: `64`

**6. 2D Classification**\
\
**7. Select 2D (INTERACTIVE)**\
Select all views that look resolvable; these classes will be used to identify particle locations on the full dataset. Try to select the following views:

![](/files/-McUyzBWI8YW_iKx-jrU)

**8. Template Picker**

* Particle Diameter: `160`

**9. Inspect Picks (INTERACTIVE)**\
Move sliders to match the following values:

* NCC score > `0.340`
* Local power > `936.000`
* Local power < `1493.000`

**10. Extract from Micrographs**

* Extraction Box Size: `360`
* Fourier crop to box size: `256`

**11. 2D Classification**

**12. Select 2D (INTERACTIVE)**

![](/files/-McUzYOB1onvVxobkcgD)

**13. Ab-Initio Reconstruction**

**14. Non-Uniform Refinement**\
The Non-Uniform Refinement should yield a 3Å resolution structure, at which point the data processing pipeline is complete.

Your tree view should look like this:

![The final tree view for the processing pipeline.](/files/-McV-S8JseQ5P9Xl9ymr)

## **Cluster Configuration**

[AWS ParallelCluster](https://aws.amazon.com/hpc/parallelcluster/) 2.10.0 was used to create an HPC cluster on which to benchmark CryoSPARC. The main requirements for cryo-EM workloads are access to large numbers of GPUs (and CPUs for some applications) as well as a high-performance file system. AWS ParallelCluster provides a simple-to-use mechanism to create a cluster that meets those requirements.

The specific cluster architecture for these benchmarks is as follows (deployed in us-east-1):

![](/files/-M_Lb3viHHPUbJy91Epp)

### **Main Node**

* EC2 instance: c5n.9xlarge
* 36 vCPUs
* 96 GB memory
* Network Bandwidth: 50 Gbps
* 100 GB local storage (EBS gp2)

### **Compute Nodes**

Multiple queues were configured to dynamically provision the following instance types:

* g4dn.16xlarge
  * Intel Cascade Lake CPU (32 vCPUs, 256GB memory)
  * 1 x NVIDIA T4 GPU
  * 1 x 900 GB NVMe local storage
* g4dn.metal
  * Intel Cascade Lake CPU (32 vCPUs, 256GB memory)
  * 8 x NVIDIA T4 GPUs
  * 2 x 900 GB NVMe local storage
* p3.2xlarge
  * Intel Broadwell CPU (8 vCPUs, 61GB memory)
  * 1 x NVIDIA Tesla V100 GPU
* p3.8xlarge
  * Intel Broadwell CPU (32 vCPUs, 244GB memory)
  * 4 x NVIDIA Tesla V100 GPUs
* p3.16xlarge
  * Intel Broadwell CPU (64 vCPUs, 488GB memory)
  * 8 x NVIDIA Tesla V100 GPUs
* p3dn.24xlarge
  * Intel Broadwell CPU (96 vCPUs, 768GB memory)
  * 8 x NVIDIA Tesla V100 GPUs
  * 2 x 900 GB NVMe local storage
* p4d.24xlarge
  * Intel Cascade Lake CPU (96 vCPUs, 1152GB memory)
  * 8 x NVIDIA A100 GPUs
  * 8 x 1 TB NVMe local storage

Further details about the above EC2 instance types can be found [here](https://aws.amazon.com/ec2/instance-types/).

### **File Systems**

* `/fsx`
  * 12 TB FSx for Lustre (2.4 GB/s throughput)
  * Used as the primary working directory for all CryoSPARC jobs
* `/shared`
  * 100 GB EBS volume mounted on cluster head node
  * Shared with compute nodes via NFS
  * Used as application installation directory
* `/scratch`
  * Local storage on compute nodes
  * Only used to test cache performance in certain CryoSPARC steps
  * Only available on specific EC2 instances (those with additional NVMe or SSD)

![Storage architecture of a p4d.24xlarge in AWS ParallelCluster](/files/-M_Lb3vjMikpBkLfSHaD)

### **Software**

* CryoSPARC v3.0.0
* AWS ParallelCluster v2.10.0
* Slurm 20.02.4
* CUDA 11.0
* NVIDIA Driver 450.80.02

## **EMPIAR 10288 Pipeline**

The pipeline used to benchmark the EMPIAR-10288 dataset is composed of the following CryoSPARC steps:

* Import Movies
* Patch Motion Correction
* Patch CTF Estimation
* Blob Picker
* Template Picker
* Extract From Micrographs
* 2D Classification
* Ab-initio Reconstruction
* Non-Uniform Refinement

### **Performance Analysis**

Each stage was run on 1, 2, 4 or 8 GPUs on each of the listed EC2 instance types (certain stages only make use of a single GPU and are noted in the results). Each pipeline step was run on the attached FSx for Lustre filesystem, and several of the steps were also run using local NVMe drives as a cache to compare performance.

The total runtime for the EMPIAR-10288 pipeline on each instance type is shown below:

![](/files/-M_Lb3vkgH7QURL0ndne)

![](/files/-M_Lb3vlJ9PGpw0PqUdh)

The p4d instance provides the best overall performance, but the analysis pipeline doesn’t make use of all 8 GPUs the entire time. Also, the cost of running the entire pipeline on a single p4d instance may not be ideal for some users:

![](/files/-M_Lb3vm_rKN4cVwylfW)

The g4dn.metal instance provides the most cost-effective option if we were to use it for the entire analysis.

An ideal approach is to match EC2 instance types to CryoSPARC pipeline stages, allowing us to make efficient use of the compute resources and keep compute costs down. The benchmarks for each step are listed below and will be used to help identify which instance types to use for which step.

![](/files/-M_Lb3vneu2DnNulE-Eb)

The movie import step does not make use of GPUs, so the performance is determined by the host CPU. The g4 and p4 instances have Intel Cascade Lake processors, the p3dn.24xlarge has a Skylake processor, and the p3 instances have Broadwell processors.

![](/files/-M_Lb3voOu02V9ObA_Tx)

![](/files/-M_Lb3vpjaVfosZ9_9cq)

The Patch Motion and Patch CTF Estimation steps show good scaling, and so an instance with 8 GPUs is recommended for these stages.

![](/files/-M_Lb3vqKWFI0TGxJrIu)

![](/files/-M_Lb3vrLM44NtZZaDKD)

The Blob Picker and Template Picker stages make use of a single GPU, and there is minimal difference in performance between GPU types. Here a low-cost, single-GPU EC2 instance is recommended (e.g. g4dn.16xlarge).

![](/files/-M_Lb3vstZC81JgiNfrU)

The Extract from Micrographs step shows scaling only up to 2 GPUs. Currently, there are no EC2 instances that offer only 2 GPUs, so a 4-GPU instance is recommended (e.g. p3.8xlarge). The performance difference is minimal across GPU architectures, so price should be the driving factor in choosing an instance for this stage.

![](/files/-M_Lb3vti21P2zlf306V)

The 2D Classification step also shows scaling only up to 2 GPUs. At present, EC2 instances are available with 1, 4, or 8 GPUs, so a 4-GPU instance is recommended (e.g. p3.8xlarge). Optionally, a lower-cost instance like the g4dn.metal (with 8 GPUs) may be used, with the expectation that scaling may be limited. The performance difference is minimal across GPU architectures, so price should be the driving factor in choosing an instance for this stage.

![](/files/-M_Lb3vunkMlKv3tOs07)

The Ab-initio Reconstruction step makes use of a single GPU, so the g4dn.16xlarge instance is the best option in terms of price and performance. The larger p3 and p4 instances offer the best performance but increase the overall cost.

![](/files/-M_Lb3vvPoQ6srl64t_Z)

The Non-Uniform Refinement stage makes use of 1 GPU and contributes the most to the total runtime. For this stage, a user should consider whether cost or time-to-solution is more important; a p4d.24xlarge will produce results 3.4x faster than a g4dn.16xlarge, but at a higher cost.

The following data shows a breakdown of each stage as a percentage of the total runtimes.![](/files/-M_Lb3vwRdzJ0oeN5QrL)![](/files/-M_Lb3vxD5iUX6vxwnUZ)![](/files/-M_Lb3vyhD7Kv6K9DY6p)![](/files/-M_Lb3vz-glbA1I1ybIC)

We can see several different aspects of a simulation to consider when choosing instance types:

* The Patch Motion Correction and Patch CTF Estimation stages scale, and so an 8-GPU instance is recommended.
* Ab-initio Reconstruction and Non-Uniform Refinement dominate the runtime as single-GPU stages.
* 2D Classification benefits somewhat from scaling, but the short run time for this stage means cost should be a deciding factor in choosing an instance.

As CryoSPARC involves an analysis pipeline, with each stage making use of compute resources in a different way, care must be given to the selection of instance type.

### **Storage Performance**

Another key component in the CryoSPARC pipeline is the filesystem. It is recommended to make use of a local, fast storage on the compute node such as a cache (e.g. NVMe). AWS EC2 instances do offer NVMe local drives, but only a subset of the instances. Generally, they are available on larger instances (e.g., p4d.24xlarge, p3dn.24xlarge, and g4dn.metal). In order to provide a cost-optimal solution and allow for smaller GPU instance types, FSx for Lustre was benchmarked to determine if it could provide the necessary performance.

Below is benchmark data for the 2D classification step (one of the steps in CryoSPARC that can make use of a local cache). It shows the percentage decrease in runtime using FSx for Lustre.

In every case except for one, FSx for Lustre provided better performance than using local storage as a cache.

![](/files/-M_Lb3w-VhOtRe1g62lh)

The main reason for this is how the local filesystem is created. Larger instances have multiple drives (e.g. the p4d.24xlarge has 8 NVMe drives). AWS ParallelCluster creates a single, logical volume from those drives, and that logical volume affects the performance. Additionally, the 2D classification step can make use of multiple GPUs, and so we see a performance loss when multiple cryoSPARC tasks are using the same filesystem (as opposed to those tasks using FSx for Lustre, a file system designed to support IO for a large number of processes).<br>

### **Cost Analysis**

Using a single EC2 instance type for an entire CryoSPARC workload is not an optimal solution. Instances like the p4d.24xlarge will provide the best performance, but with no consideration to cost. General guidelines were given above as to what instances to pick for what stage. Below is an example of how a user may implement those recommendations, along with the total cost (including storage), and runtime. Prices are listed in USD for the us-east-1 region.

|                                                                            | **Configuration 1** |                | **Configuration 2**                                              |                   |                |                                                                 |
| -------------------------------------------------------------------------- | ------------------- | -------------- | ---------------------------------------------------------------- | ----------------- | -------------- | --------------------------------------------------------------- |
| **Pipeline Stage**                                                         | **Instance Type**   | **Cost (USD)** | **Runtime (Min.)**                                               | **Instance Type** | **Cost (USD)** | **Runtime (Min.)**                                              |
| **Patch Motion Correction**                                                | p4d.24xlarge        | $12.25         | 22.4                                                             | g4dn.metal        | $3.97          | 30.5                                                            |
| **Patch CTF Estimation**                                                   | p4d.24xlarge        | $7.46          | 13.7                                                             | g4dn.metal        | $1.44          | 11.1                                                            |
| **Blob picker**                                                            | g4dn.16xlarge       | $0.72          | 9.9                                                              | g4dn.16xlarge     | $0.72          | 9.9                                                             |
| **Template Picker**                                                        | g4dn.16xlarge       | $1.77          | 24.4                                                             | g4dn.16xlarge     | $1.77          | 24.4                                                            |
| **Micrograph Extraction**                                                  | p3.8xlarge          | $1.65          | 8.25                                                             | g4dn.16xlarge     | $0.90          | 12.5                                                            |
| **2D Classification**                                                      | g4dn.metal          | $1.76          | 13.5                                                             | g4dn.metal        | $1.76          | 13.5                                                            |
| **Ab-initio Reconstruction**                                               | g4dn.16xlarge       | $2.56          | 35.3                                                             | p3.16xlarge       | $6.83          | 23.3                                                            |
| **Non-Uniform Refinement**                                                 | g4dn.16xlarge       | $9.43          | 130                                                              | p4d.24xlarge      | $33.66         | 82.5                                                            |
| **Compute Cost**                                                           |                     | $37.60         |                                                                  |                   | $51.05         |                                                                 |
| <p><strong>Lustre Cost</strong></p><p><strong>(12 TB Scratch)</strong></p> |                     | $9.52          |                                                                  |                   | $8.03          |                                                                 |
| **Total**                                                                  |                     | **$47.12**     | <p><strong>257.55</strong></p><p><strong>(4.29 hrs)</strong></p> |                   | **$59.08**     | <p><strong>206.8</strong></p><p><strong>(3.45 hrs)</strong></p> |

For comparison, here are costs and runtimes if we ran the entire benchmark on a single EC2 instance type:

* p4d.24xlarge
  * **2.79 hours**
  * **$91.40**
* g4dn.metal
  * **4.8 hours**
  * **$37.60**
* p3.16xlarge
  * **3.4 hours**
  * **$83.00**

## **Conclusion**

The benchmark results presented demonstrate how AWS can be used to accelerate cryo-EM workloads. By making use of AWS ParallelCluster, users can easily create an HPC cluster with a range of GPU instance types, allowing them to best match compute resources with the requirements of each analysis step.

Benchmark data presented here was done with a single user; in practice, it’s likely multiple users will use the cluster for analysis. Further cost optimization can be achieved by running multiple jobs on a larger instance. For example, while a p4d.24xlarge (with 8 A100 GPUs) may have a higher cost, running multiple, single-GPU stages like Non-Uniform Refinement at the same time will help amortize the higher cost of the p4d instance.

Summarized below are the key points a user should consider when creating a cryoEM cluster.

* Amazon FSx for Lustre provides a high-performance file system that can meet the requirements of a cryo-EM analysis pipeline. This also allows flexibility in choosing instance types (e.g. those that lack local, fast NVMe storage).
* A range of GPU instances should be employed. AWS ParallelCluster can be used to easily create different queues with different EC2 instances for this purpose.
* p3dn.24xlarge instances are not recommended. They can provide excellent performance, but the p4d.24xlarge instance is priced very closely to the p3dn.24xlarge. The time-to-solution for the p4d.24xlarge is fast enough that in most cases a processing stage will use less compute time and, thus, cost less.
* g4dn instances will likely make up the bulk of the compute resources. They provide performance at an excellent price point.


# Troubleshooting

Overview of common issues and advice on how to resolve them.

{% hint style="info" %}
Unless otherwise noted:

* Log in to the workstation or remote node where `cryosparc_master` is installed.
* Use the same non-root UNIX user account that runs the CryoSPARC process and was used to install CryoSPARC.
* Run all commands on this page in a terminal running `bash`

\
In v4.0+, you can download error reporting information from within the application. For more details, see: [Guide: Download Error Reports](/setup-configuration-and-management/software-system-guides/guide-download-error-reports)
{% endhint %}

## Common Issues

### Cannot download or install CryoSPARC

Problems with the [**installation**](/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc) steps are indicated with the some of the following error messages:

> "Couldn't connect to host"\
> "Could not resolve host"\
> {"success": false}\
> "tar: This does not look like a tar archive"\
> "Version mismatch! Worker and master versions are not the same. Please update."\
> "An unexpected error has occurred."

**Steps**

1. If you have [Conda](https://docs.conda.io/projects/conda/en/latest/index.html) installed, [deactivate any active environments](https://docs.conda.io/projects/conda/en/latest/user-guide/tasks/manage-environments.html#deactivating-an-environment).
2. Check that your `LICENSE_ID` environment variable is set correctly with this command

   ```bash
   echo $LICENSE_ID
   ```

   Ensure the output exactly matches the CryoSPARC License ID issued to you over email.
3. Check your machine's connection to CryoSPARC's license servers at get.cryosparc.com with this `curl` command:

   ```bash
   curl https://get.cryosparc.com/checklicenseexists/$LICENSE_ID
   ```

   \
   You should see the message `{"success": true}`. If instead you see `{"success": false}`, your license is not valid, so please check it has been entered correctly.\
   \
   If you see an error message like **"Couldn't connect to host"** or **"Could not resolve host"** check your Internet connection, firewall or ensure your IT department allows access to the `get.cryosparc.com` license server domain.

### Cannot update CryoSPARC

**Steps**

* [Reinstall CryoSPARC](/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc), following the steps from the ["Cannot download CryoSPARC"](#cannot-download-cryosparc) section

### CryoSPARC does not start or encounters error on startup

This can happen following a fresh install or recent update.

**Steps**

1. In a command line, run `cryosparcm status`
2. Check that the output looks like this

   ```

   CryoSPARC System master node installed at
   /home/cryosparcuser/cryosparc/cryosparc_master
   Current CryoSPARC version: v5.0.0
   --------------------------------------------------------------------------------

   CryoSPARC process status:

   api:0                            RUNNING   pid 2439162, uptime 27 days, 19:52:14
   api:1                            RUNNING   pid 2439163, uptime 27 days, 19:52:14
   api:2                            RUNNING   pid 2439164, uptime 27 days, 19:52:14
   app                              RUNNING   pid 2443242, uptime 27 days, 19:52:04
   app_api                          RUNNING   pid 2443078, uptime 27 days, 19:52:06
   cache                            RUNNING   pid 2435460, uptime 27 days, 19:52:36
   command_vis                      RUNNING   pid 2440324, uptime 27 days, 19:52:08
   database                         RUNNING   pid 2435129, uptime 27 days, 19:52:38
   scheduler                        RUNNING   pid 2439743, uptime 27 days, 19:52:10
   search                           RUNNING   pid 2440388, uptime 27 days, 19:52:07

   --------------------------------------------------------------------------------
   ✓ License is valid
   --------------------------------------------------------------------------------

   export CRYOSPARC_LICENSE_ID="xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
   export CRYOSPARC_MASTER_HOSTNAME="localhost"
   export CRYOSPARC_DB_PATH="/home/cryosparcuser/cryosparc/cryosparc_database"
   export CRYOSPARC_BASE_PORT=61000
   export CRYOSPARC_INSECURE=false
   export CRYOSPARC_CLICK_WRAP=true
   ```
3. Check that all items under "CryoSPARC process status" that *do not* end in `_dev` or `_legacy` are `RUNNING`. If any are not, run `cryosparcm restart`
4. If any non-`_dev` /non-`_legacy` components have a status other than `RUNNING` (such as `STOPPED` or `EXITED`), check their log for errors. For example, this command checks for errors on the `database` process:

   ```bash
   cryosparcm log database
   ```

   (Press `control C`, then `q` to stop logging)
5. If the web interface is inaccessible, check firewall settings to ensure [CryoSPARC's base port number](/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc#optional-re-load-your-bashrc-this-will-allow-you-to-run-the-cryosparcm-management-script-from-anywhere-in-the-system) is exposed for network access

Any error messages here could indicate specific configuration issues and may require [**re-installing CryoSPARC**](/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc).

If at any point you see `No command 'cryosparcm' found` or `command not found: cryosparcm`:

1. Check that you are on the master node or workstation where `cryosparc_master` is installed
2. Run `echo $PATH` and check that it contains `<installation directory>/cryosparc_master/bin`

   ```bash
   $ echo $PATH
   /home/cryosparcuser/cryosparc/cryosparc_master/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
   ```
3. [Reinstall CryoSPARC](/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc) if the above did not restore the the `cryosparcm` command

### 'User not found' error when attempting to log in

This error message occurs if the email address field does not match any existing users in your CryoSPARC instance. Use the CryoSPARC command-line to verify the details of your user account and change the email address or password if needed.

1. Run the following command in your terminal: `cryosparcm listusers`
2. If an email address is incorrect (e.g., mispelled or with an extra space at the beginning or end), modify it in the database. Run the following commands:

   * Log into the MongoDB shell: `cryosparcm mongo`
   * Once in the MongoDB shell, enter the following (replace the incorrect/correct email):

   ```javascript
   db.users.update(
     { 'emails.0.address': 'incorrect@domain.eud' }, 
     { $set: { 'emails.0.address': 'correct@domain.edu' } 
   })
   ```

   * Exit the MongoDB shell with `exit`
3. If you don't remember your password, reset it with the following command (replace with your email address and new password):

   ```bash
   cryosparcm user resetpassword --email "<email address>"
   ```

   Enter the new password when prompted.

### Recover from incomplete CryoSPARC shutdown

An incomplete shutdown of CryoSPARC is likely to interfere with subsequent attempts to start CryoSPARC and/or CryoSPARC software updates. Incomplete shutdowns can occur for various reasons, including, but not limited to:

* unclean shutdown of the computer that runs `cryosparc_master` processes
* failed coordination of services by `cryosparc_master`'s `supervisord` process

Follow this sequence to ensure a complete shutdown of CryoSPARC

#### 1. Basic shutdown

For CryoSPARC instances that were *not* configured as a systemd service, run the command

```
cryosparcm stop
```

{% hint style="warning" %}
Do not use

```
cryosparcm stop
```

for CryoSPARC instances that are controlled by systemd. For such instances, use the appropriate `systemctl stop` command.
{% endhint %}

#### 2. Find and, if necessary, terminate "zombie" processes

Confirm that the *basic shutdown* did not "leave behind" any CryoSPARC-related processes. *If* the *basic shutdown* was successful, a suitable `ps` command should not show any processes for the *CryoSPARC instance in question*, but processes may be shown if

* a glitch occurred during the *basic shutdown* or
* the computer hosts multiple CryoSPARC instances.

To illustrate what kind of processes one might encounter, here is an example command and its output for a running CryoSPARC v5.0 instance:

```shell
$ ps -weo pid,ppid,start,cmd | grep -e cryosparc -e mongo | grep -v grep
2017484       1   Jun 30 python3.12 supervisord -c /home/cryosparcuser/cryosparc/cryosparc_master/config/supervisord.conf
2024251 2017484   Jun 30 mongod --dbpath /home/cryosparcuser/cryosparc/cryosparc_database --port 61001 --oplogSize 64 --replSet meteor --wiredTigerCacheSizeGB 4 --bind_ip_all --networkMessageCompressors snappy
2026655 2017484   Jun 30 python3.12 uvicorn api.main:app --fd 0 --log-config /home/cryosparcuser/cryosparc/cryosparc_master/config/log.conf
2026657 2017484   Jun 30 python3.12 uvicorn api.main:app --fd 0 --log-config /home/cryosparcuser/cryosparc/cryosparc_master/config/log.conf
2026658 2017484   Jun 30 python3.12 uvicorn api.main:app --fd 0 --log-config /home/cryosparcuser/cryosparc/cryosparc_master/config/log.conf
2029262 2017484   Jun 30 python flask --app command.command_vis:start() run -h 0.0.0.0 -p 61003 --with-threads
```

{% hint style="info" %}
This is a *simple* example. More complex configurations, such a host with multiple active CryoSPARC instances, may require different `ps` options and/or `grep` patterns.
{% endhint %}

{% hint style="warning" %}
The `ps` output may include processes that belong to non-CryoSPARC applications or to CryoSPARC instances other than the CryoSPARC instance that you wish to shutdown. Parent process identifiers and port numbers in the listed commands can help in attributing processes to a common parent `supervisord` process. Carefully confirm the purpose and identity of any process before termination.
{% endhint %}

For the example above, it should be sufficient to `kill` the `supervisord` process using the process identifier shown by the `ps` command

```
kill 2017484
```

and wait a few seconds for the `supervisord` process' children to be terminated automatically.

{% hint style="danger" %}
[Never](https://www.mongodb.com/docs/v3.6/tutorial/manage-mongodb-processes/#sigkill) use the `kill -9` option for `mongod` processes.
{% endhint %}

Finally, using another `ps` command with suitable options, re-confirm that all relevant processes have in fact been terminated

#### 3. Only under certain circumstances, delete "orphaned" mongodb socket file

{% hint style="info" %}
Socket files should be deleted only under specific circumstances, subject to precautions given below.
{% endhint %}

An intact CryoSPARC instance manages the creation and deletion of the socket file for `mongod`, like

```
/tmp/mongodb-61001.sock
```

Filenames differ between CryoSPARC instances, for example based on `$CRYOSPARC_DB_PORT`.

{% hint style="danger" %}
Never delete socket files before confirming that associated processes have been terminated as described in the previous step.
{% endhint %}

{% hint style="warning" %}
The computer may store socket files that belong to non-CryoSPARC applications or to CryoSPARC instances other than the CryoSPARC instance that you wish to shutdown. Such socket files may have names similar to the files you wish to delete. Carefully confirm the purpose and identity of each file before any deletion.
{% endhint %}

### License error or license not found

Follow the steps in this section when you see error messages that look like this:

> "License is invalid."\
> "License signature invalid."\
> "Could not find local license file. Please re-establish your connection to the license servers."\
> "Local license file is expired. Please re-establish your connection to the license servers."\
> "Token is invalid. Another CryoSPARC instance is running with the same license ID."

**Steps**

Run `cryosparcm licensestatus`. This should result in "License is valid". If you see this error:

```
ServerError: Authentication failed
```

Your license ID is not entered configured correctly. Check the `CRYOSPARC_LICENSE_ID` entry in `cryosparc_master/config.sh` .

If you see this error:

```
WARNING: Could NOT verify active license
```

This includes a list of checks, the last of which will indicate a failure. Depending on which check failed, do one of the following:

* Ensure you entered the license correctly during the [installation step](/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc).
* Check your Internet connection
* Check your machine's connection to CryoSPARC's license servers at get.cryosparc.com with this `curl` command (substitute `<license>` with your unique license ID):

```bash
curl https://get.cryosparc.com/checklicenseexists/<license>
```

Look for the message message `{"success": true}`

If instead you see `{"success": false}`, your license is not valid so please check it has been entered correctly.

If you see an error message like **"Couldn't connect to host"** or **"Could not resolve host"** check your Internet connection, firewall or ensure your IT department has the `get.cryosparc.com` license server domain whitelisted.

If you see a license ID conflict such as

> "Another cryoSPARC instance is running with the same license ID."

Follow the [*complete* shutdown procedure](#recover-from-incomplete-cryosparc-shutdown) before running

```bash
cryosparcm start
```

### Cannot queue or run job

Follow these steps when the CryoSPARC web interface is up and running normally and jobs may be created but do not run. These error messages may indicate this issue:

> "list index out of range"\
> "Could not resolve hostname ... Name or service not known"

A job that never changes from `Queued` or `Launched` status may also indicate this.

**Steps**

1. Ensure at least one worker is connected to the master. See the [Installation](/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc) page for details. Visit **Manage > Resources** to see what lanes are available
2. Check that all non-development CryoSPARC processes are running with the [`cryosparcm status`](/setup-configuration-and-management/management-and-monitoring-4.7/cryosparcw-4.7#cryosparcm-status) command
3. (For master/worker setups) check that SSH is configured between the master and worker machines.
4. Check the log for the `api` and `scheduler` services to find any application error messages

   ```bash
   cryosparcm log api
   cryosparcm log scheduler
   ```

   (for each command, press `Control + C` on the keyboard to allow scrolling up. Press  `q` to exit when finished)
5. If applicable, check that the [cluster submission script](/setup-configuration-and-management/management-and-monitoring-4.7/cryosparcw-4.7#cryosparcm-cluster) is correct
6. Stop CryoSPARC completely:

   ```bash
   cryosparcm stop
   ```
7. Force-reinstall [master](/setup-configuration-and-management/management-and-monitoring-4.7/cryosparcm-4.7#cryosparcm-forcedeps) and [worker](/setup-configuration-and-management/management-and-monitoring-4.7/cryosparcw-4.7#cryosparcw-forcedeps) dependencies. This can help when a worker was not correctly installed.\
   \
   On the workstation or master run:

   ```
   cryosparcm deps --force
   ```

   \
   On workers run:

   ```
   cryosparcw deps --force
   ```
8. Restart CryoSPARC:

   ```
   cryosparcm start
   ```
9. Clear the job and re-run it.

### Job stuck in launched status

This indicates that CryoSPARC started the job process but the job encountered an internal error immediately after.

#### Steps

Check the job's **Event Log** for errors. If there are none, check the standard out log either from the interface under **Job > Metadata > Job Log**\ <img src="/files/bVd9ptAudXgWjVelfBXG" alt="Job Log in the job&#x27;s metadata tab" data-size="original">

or from the command line with

```
cryosparcm job log PX JY
```

Substituting `PX` and `JY` for the project and job IDs, respectively.

Typically the errors here occur when the worker process cannot connect back to the master. Ensure there is a stable network connection between all machines involved. Ensure CryoSPARC was installed correctly and [re-install](/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc) if necessary.

### Job runs but ends unexpectedly with status "Failed"

When a job fails, its job card in the interface turns red and the bottom of the job log includes an error message with the text `Traceback (most recent call last)`

Common failure reasons include:

* Invalid or unspecified input slots
* Invalid or unspecified required parameters, including file/folder paths
* Incorrectly set up GPU (e.g., running a job on a node without enough GPUs or [CUDA](https://nvidia.com/cuda) drivers not installed)
* Another process taking up memory on a GPU
* Cache not set up correctly for a worker
* Lost connection to `cryosparc_master`

![Example of failed jobs in a CryoSPARC workspace](/files/LmQ1aFkg5GN4W6ny56HO)

![Example of an error log entry at the bottom of a Failed job](/files/-M7DHKGkP2POQy3RAIqc)

Common job failure error messages:

> "AssertionError: Child process with PID ... has terminated unexpectedly!"\
> Job is unresponsive - no heartbeat received in 30 seconds.

Common error messages that indicate incorrectly configured GPU drivers:

> "cuInit failed: unknown error"\
> "no CUDA-capable device is detected"\
> "cuMemHostAlloc failed: OS call failed or operation not supported on this OS"\
> "cuCtxCreate failed: invalid device ordinal\
> kernel.cu ... error: identifier "\_\_shfl\_down\_sync" is undefined

Common error messages that indicate not enough GPU memory:

> "cuMemAlloc failed: out of memory"\
> "cuArrayCreate failed: out of memory"\
> "cufftAllocFailed"

**Steps**

1. Ensure a [supported version](https://guide.cryosparc.com/setup-configuration-and-management/pages/-M7DHIJrIWYpsjbmcVFX#1.-cuda-toolkit) of the [CUDA](https://developer.nvidia.com/cuda-downloads) toolkit is installed and running on the workstation or each worker node
2. Check the GPU configuration on the workstation or node where the job runs on. Log into that machine and navigate to the CryoSPARC installation directory. Run the `cryosparcw gpulist` command:

   ```bash
   cd /path/to/cryosparc_worker
   bin/cryosparcw gpulist
   ```
3. Run nvidia-smi to check that no other processes are using GPU memory. CryoSPARC-related process appear with process name "python"

   ![Example output of the nvidia-smi command, showing CUDA 10.2 and a CryoSPARC python process using \~2GB on GPU 0](/files/-M7DHKGlJRjhcsfBF6G0)

   If you don't recognize the processes using GPU memory, run `kill <PID>`, substituting `<PID>` with the value under the Processes PID column<br>
4. Check the Event log: Select the Job card in the CryoSPARC interface and press the Spacebar on your keyboard to see the log. Scroll down to the bottom and look for the failure reasons in red\
   ![](/files/FcO4VSWFDMvu80zDJIZM)
5. Clear the job and select the "Build" status badge on the job card to enter the Job Builder
6. If the job failed with a GPU-related error and multiple GPUs are available, try running the job on a different GPU. Press Queue, switch to the "Run on specific GPU" Queue type and select one or more GPUs\
   ![](/files/UOajNEF9fI1TxvveperK)
7. If the job failed with `AssertionError: Non-optional inputs from the following input groups and their slots are not connected` then clear the job, enter the Builder and expand any input groups connected to the job. Missing required slots appear with text "Empty" and "Required"\
   ![](/files/T4mFZJLjupu04RbGM4UX)
8. Check the job parameters: To learn about setting specific parameters, hover over or touch the parameter names in the Job Builder to see a description of what they do

   ![On-hover description of the "Negative Stain Data" parameter for the "Import Movies" job](/files/I88EPfGWclpaFAjo9omf)
9. Find the target job type in this guide's Job Reference for more detailed descriptions of expected input slots and parameters: [All Job Types in CryoSPARC](/processing-data/all-job-types-in-cryosparc).
10. Reduce the box-size of extracted particles. Some jobs need to fit thousands of particles in GPU memory at a time, and larger box sizes exceed GPU memory limits. Either extract with a smaller box size or with the [Downsample Particles job](/processing-data/all-job-types-in-cryosparc/extraction/job-downsample-particles).
11. Look for extended error information with the [`cryosparcm job log` command](/setup-configuration-and-management/management-and-monitoring-v5.0/cryosparcm-reference-v5.0#cryosparcm-job-log) (press `Ctrl + C` on the keyboard to exit when finished)
12. Check the network connection from the worker machine to the master
13. On occasion, a job fails due to an error in the CryoSPARC code (bug). The CryoSPARC team regularly releases updates and patches with bug fixes. [Check for](https://cryosparc.com/updates) and [install the latest update](/setup-configuration-and-management/software-updates) or [patch](/setup-configuration-and-management/software-updates#patches).\
    \
    If you find a new bug, see the [Additional Help](#additional-help) section for advice.

### Job stuck or taking a very long time

Due to their large sizes, cryo-EM datasets can take a long time to process with sub-optimal hardware or parameters. Here are some facilities that CryoSPARC provides for increasing speed/performance.

* [Connect workers with SSD cache enabled](/setup-configuration-and-management/management-and-monitoring-4.7/cryosparcw-4.7#cryosparcw-connect). This speeds up processing for extracted particles during 2D Classification, *ab-initio* reconstruction, refinement and more. Ensure the "**Cache particle images on SSD**" parameter is enabled under **"Compute settings"** for the target particle-processing job
* Some jobs (motion correction, ctf estimation, 2D classification) can run on multiple GPUs. If your hardware supports it, increase the number of GPUs to parallelize over

![2D Classification jobs support particle SSD caching and parallelizing over multiple GPUs.](/files/X1eOHf2n0TByyi0S1quI)

* Extracted particles with large box sizes (relative to their pixel size) take a long time to process. Consider Fourier-cropping (or "binning") the extracted particle blobs with the Downsample Particles job

{% content-ref url="/pages/17vxAkUji9lhwxmzL6P5" %}
[Job: Downsample Particles](/processing-data/all-job-types-in-cryosparc/extraction/job-downsample-particles)
{% endcontent-ref %}

* Minimize the number of processes using system resources on the workstation or worker nodes
* Check for zombie processes on worker machines. The process is similar to the steps under **"Another CryoSPARC instance is running with the same license ID"** under the [License error or license not found](#license-error-or-license-not-found) section
* Cancel the job, clear and re-queue

### GPU Issues

#### `cudaErrorInsufficientDriver` or `CUDA_ERROR_UNSUPPORTED_PTX_VERSION`

A job fails with errors similar to the following when the Nvidia driver is out-of-date or incompatible with the target GPU or CUDA Toolkit Version that ships with CryoSPARC:

```python
File "/u/cryosparc/cryosparc_worker/cryosparc_compute/jobs/runcommon.py", line 1711, in run_with_except_hook 
    run_old(*args, **kw) 
File "cryosparc_worker/cryosparc_compute/engine/cuda_core.py", line 129, in cryosparc_compute.engine.cuda_core.GPUThread.run 
File "cryosparc_worker/cryosparc_compute/engine/cuda_core.py", line 130, in cryosparc_compute.engine.cuda_core.GPUThread.run 
File "cryosparc_worker/cryosparc_compute/engine/engine.py", line 997, in cryosparc_compute.engine.engine.process.work 
File "cryosparc_worker/cryosparc_compute/engine/engine.py", line 106, in cryosparc_compute.engine.engine.EngineThread.load_image_data_gpu 
File "cryosparc_worker/cryosparc_compute/engine/gfourier.py", line 33, in cryosparc_compute.engine.gfourier.fft2_on_gpu_inplace 
File "/u/cryosparc/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.7/site-packages/skcuda/fft.py", line 102, in __init__ 
    capability = misc.get_compute_capability(misc.get_current_device()) 
File "/u/cryosparc/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.7/site-packages/skcuda/misc.py", line 254, in get_current_device 
    return drv.Device(cuda.cudaGetDevice()) 
File "/u/cryosparc/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.7/site-packages/skcuda/cudart.py", line 767, in cudaGetDevice 
    cudaCheckStatus(status) 
File "/u/cryosparc/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.7/site-packages/skcuda/cudart.py", line 565, in cudaCheckStatus 
    raise e 
skcuda.cudart.cudaErrorInsufficientDriver
```

```python
Traceback (most recent call last):
    driver.cuLinkAddData(self.handle, input_ptx, ptx, len(ptx),
  File "/home/cryosparcuser/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.8/site-packages/numba/cuda/cudadrv/driver.py", line 352, in safe_cuda_api_call
    return self._check_cuda_python_error(fname, libfn(*args))
  File "/home/cryosparcuser/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.8/site-packages/numba/cuda/cudadrv/driver.py", line 412, in _check_cuda_python_error
    raise CudaAPIError(retcode, msg)
numba.cuda.cudadrv.driver.CudaAPIError: [CUresult.CUDA_ERROR_UNSUPPORTED_PTX_VERSION] Call to cuLinkAddData results in CUDA_ERROR_UNSUPPORTED_PTX_VERSION

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "cryosparc_master/cryosparc_compute/run.py", line 96, in cryosparc_master.cryosparc_compute.run.main
  File "/home/cryosparcuser/cryosparc_worker/cryosparc_compute/jobs/instance_testing/run.py", line 174, in run_gpu_job
    func = mod.get_function("add")
  File "/home/cryosparcuser/cryosparc_worker/cryosparc_compute/gpu/compiler.py", line 256, in get_function
    cufunc = self.get_module().get_function(name)
  File "/home/cryosparcuser/cryosparc_worker/cryosparc_compute/gpu/compiler.py", line 212, in get_module
    linker.add_cu(s, k)
  File "/home/cryosparcuser/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.8/site-packages/numba/cuda/cudadrv/driver.py", line 3022, in add_cu
    self.add_ptx(program.ptx, ptx_name)
  File "/home/cryosparcuser/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.8/site-packages/numba/cuda/cudadrv/driver.py", line 3010, in add_ptx
    raise LinkerError("%s\n%s" % (e, self.error_log))
numba.cuda.cudadrv.driver.LinkerError: [CUresult.CUDA_ERROR_UNSUPPORTED_PTX_VERSION] Call to cuLinkAddData results in CUDA_ERROR_UNSUPPORTED_PTX_VERSION
ptxas application ptx input, line 9; fatal   : Unsupported .version 7.8; current version is '7.3'
```

To fix this, update the Nvidia driver to the minimum driver version noted in [Installation Prerequistes](/setup-configuration-and-management/cryosparc-installation-prerequisites). Please follow instructions specific to the worker's Linux distribution to install the Nvidia driver. The latest Nvidia driver is available to download on [Nvidia's website](https://www.nvidia.com/Download/index.aspx).

[Related Discussion Forum post where a user encountered this error.](https://discuss.cryosparc.com/t/cryosparc-tries-to-run-worker-in-old-directory-cannot-find/5632/5)

#### **`undefined symbol: _ZSt28__throw_bad_array_new_lengthv`**

A job may fail with the following error when running GPU jobs (Patch Motion Correction, 2D Classification) with CryoSPARC v3.3.2 or v3.4.0 on Ubuntu 22+.

```python
Traceback (most recent call last):
  File "cryosparc_worker/cryosparc_compute/run.py", line 72, in cryosparc_compute.run.main
  File "/home/cryosparc/cryosparc/cryosparc_worker/cryosparc_compute/jobs/jobregister.py", line 371, in get_run_function
    runmod = importlib.import_module(".."+modname, __name__)
  File "/home/cryosparc/cryosparc/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.7/importlib/__init__.py", line 127, in import_module
    return _bootstrap._gcd_import(name[level:], package, level)
  File "<frozen importlib._bootstrap>", line 1006, in _gcd_import
  File "<frozen importlib._bootstrap>", line 983, in _find_and_load
  File "<frozen importlib._bootstrap>", line 967, in _find_and_load_unlocked
  File "<frozen importlib._bootstrap>", line 677, in _load_unlocked
  File "<frozen importlib._bootstrap_external>", line 1050, in exec_module
  File "<frozen importlib._bootstrap>", line 219, in _call_with_frames_removed
  File "cryosparc_worker/cryosparc_compute/jobs/class2D/run.py", line 13, in init cryosparc_compute.jobs.class2D.run
  File "/home/cryosparc/cryosparc/cryosparc_worker/cryosparc_compute/engine/__init__.py", line 8, in <module>
    from .engine import *  # noqa
  File "cryosparc_worker/cryosparc_compute/engine/engine.py", line 9, in init cryosparc_compute.engine.engine
  File "cryosparc_worker/cryosparc_compute/engine/cuda_core.py", line 4, in init cryosparc_compute.engine.cuda_core
  File "/home/cryosparc/cryosparc/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.7/site-packages/pycuda/driver.py", line 62, in <module>
    from pycuda._driver import *  # noqa
ImportError: /home/cryosparc/cryosparc/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.7/site-packages/pycuda/_driver.cpython-37m-x86_64-linux-gnu.so: undefined symbol: _ZSt28__throw_bad_array_new_lengthvytho
```

To fix, set environment variables `CFLAGS="-static-libstdc++"` and `CXXFLAGS="-static-libstdc++"` to the environment before recompiling the PyCUDA module with `cryosparcw newcuda`:

```bash
cd /path/to/cryosparc_worker
export CFLAGS="-static-libstdc++"
export CXXFLAGS="-static-libstdc++"
bin/cryosparcw newcuda ~/cryosparc/cuda-11.8.0
```

### SSD Cache Issues

In some circumstances, jobs with "Cache particle images on SSD" enabled will not complete with one of the following errors in the job log:

> SSD cache needs additional $$x$$ B but drive can only be filled up to $$y$$ B

> SSD cache needs additional $$x$$ B but drive has $$y$$ B free. CryoSPARC can only free a maximum of $$z$$ B. This may indicate that programs other than CryoSPARC are using the SSD.

> Cannot allocate space on the SSD to cache /path/to/particles.mrc; other non-CryoSPARC processes are using the cache.

> Cannot finish transfer; other programs may be accessing the SSD.

> FileNotFoundError: \[Errno 2] No such file or directory: '…/…/store-v2/...' → '/scratch/instance\_cryosparc:60001/links/...’

You may also observe some of the following symptoms:

* Slower than expected SSD copy speeds
* Jobs that use the SSD cache take much longer than they should
* Jobs spend a very long time waiting on cache files locked by other jobs or waiting for additional space to free up

These issues may occur in long-active CryoSPARC instances with many projects and files on the SSD, and many jobs or other non-CryoSPARC processes running simultaneously. There are several strategies to address these or reduce their likelihood.

#### Option 1: Check SSD health

SSDs naturally degrade over time, the likelihood of failure increasing with heavy usage. Use a tool such as [`smartctl`](https://www.smartmontools.org/browser/trunk/smartmontools/smartctl.8.in) to check the SSD. If enough errors have accumulated, the SSD may have to be replaced.

CryoSPARC automatically removes files from the SSD that have not been accessed in a while (more than 30 days by default) each time the SSD cache system runs. If the SSD is very heavily used in particle processing jobs or by other external tools, leaving more free space available may extend its lifetime. This is possible with one or both of these strategies:

* [Reconnect the worker](/setup-configuration-and-management/management-and-monitoring-4.7/cryosparcw-4.7#cryosparcw-connect-less-than-options-greater-than) with the `--ssdreserve` flag set (default 10GB or 10000MB) to ensure CryoSPARC always leaves the given amount of free space on the SSD (will clean out old files to stay above the threshold)
* Set the [`CRYOSPARC_SSD_CACHE_LIFETIME_DAYS` environment variable](https://guide.cryosparc.com/setup-configuration-and-management/pages/P3sVaHfkUFPsou1J5FCe#cryosparc_master-config.sh) in `cryosparc_master/config.sh` to clean up unused files on the SSD more frequently. The default value is `30` days

#### Option 2: Ensure no other programs are using the CryoSPARC SSD cache path

CryoSPARC assumes that it has exclusive access to the SSD cache *path*. If other programs are accessing the cache *path* at the same time as CryoSPARC, the cache step may fail with a file system-related or OS-level error.

{% hint style="info" %}
Sharing a cache *device* with other applications, *including other CryoSPARC instances*, is not recommended:

* Performance may suffer from due to shared bandwidth usage.
* The given CryoSPARC instance's caching may be disrupted because the CryoSPARC instance does not prevent use of free space within its own quota by other applications or CryoSPARC instances.
  {% endhint %}

#### **Option 3: Increase or Reduce the SSD Quota**

The **SSD Quota** is the maximum amount of SSD space that CryoSPARC jobs use on a connected worker. CryoSPARC will never use more than this amount of SSD space. A job that requires caching will delete enough older cache files to ensure the quota is not exceeded.

* If the quota is significantly lower than the total size of the SSD and jobs frequently wait for cache space to free up, **consider increasing the quota**.
* If other programs regularly share the SSD with CryoSPARC and jobs fail because the cache drive runs out of space, **consider decreasing the quota** to reduce the likelihood of this.

Change the quota by [re-running the `cryosparcw connect` command](/setup-configuration-and-management/management-and-monitoring-4.7/cryosparcw-4.7#cryosparcw-connect-less-than-options-greater-than) with `--ssdquota <amount in MB>` and `--update` arguments.

#### Option 4: Increase or Reduce the SSD Reserve

The **SSD Reserve** is the minimum amount of free space that CryoSPARC leaves on the SSD, also considering potential use of the the cache *device* by other applications even if CryoSPARC has not reached its SSD quota. CryoSPARC jobs will delete, if possible, unused files from its cache until this much space is free again. CryoSPARC jobs will not copy files to the SSD until at least this amount of space is free. This setting is intended to preserve SSD health.

* If CryoSPARC reports slower cache write speeds than expected, or jobs that use the SSD cache take longer than they should, **consider increasing the reserve**.
* If jobs frequently wait for cache space to free up, **consider decreasing the reserve**.

Change the reserve by [re-running the `cryosparcw connect` command](/setup-configuration-and-management/management-and-monitoring-4.7/cryosparcw-4.7#cryosparcw-connect-less-than-options-greater-than) with `--ssdreserve <amount in MB>` (default `10000`) and `--update` arguments.

#### Option 5: Change the SSD cache locking strategy

{% hint style="info" %}
Applies to SSD caches that run on network file systems such as NFS, GPFS, BeeGFS or Lustre. Note that we strongly recommend local SSD caches for best performance.
{% endhint %}

By default, CryoSPARC jobs use a file-system level POSIX lock to ensure mutual exclusion between multiple jobs that access the SSD cache simultaneously. These locks can be unreliable on network file systems such as NFS, GPFS, BeeGFS or Lustre. This may lead to errors during the SSD caching step, particularly during copy or symbolic link operations. Error messages with the following format are key symptoms of this:

> FileNotFoundError: \[Errno 2] No such file or directory: '…/…/store-v2/...' → '/scratch/instance\_cryosparc:61001/links/...’

Add the following line to [`cryosparc_worker/config.sh`](https://guide.cryosparc.com/setup-configuration-and-management/pages/P3sVaHfkUFPsou1J5FCe#cryosparc_worker-config.sh) to use the CryoSPARC master to broker cache access instead:

```bash
export CRYOSPARC_CACHE_LOCK_STRATEGY="master"
```

#### Option 6: Fix database inconsistencies (CryoSPARC ≤4.4)

{% hint style="info" %}
These instructions apply to CryoSPARC versions v4.4 or older. v4.5+ uses a new cache system that does not require the database.
{% endhint %}

`cryosparc_master` uses its MongoDB database to coordinate SSD caching between multiple workers running in parallel. If a job fails unexpectedly during the SSD caching step, this could lead to database inconsistencies which prevent other jobs from proceeding.

To address these, try the following steps:

1. Ensure no jobs are running in CryoSPARC
2. In a terminal, enter `cryosparcm mongo` to enter the interactive database prompt
3. Enter the following command to check how many records are in an inconsistent state:

   <pre class="language-javascript"><code class="lang-javascript"><strong>db.cache_files.find({status: {$nin: ['hit', 'miss']}}).length()
   </strong></code></pre>
4. If the result is not `0` (zero), enter the following command to fix them

   ```javascript
   db.cache_files.updateMany({status: {$nin: ['hit', 'miss']}}, {$set: {status: 'miss'}})
   ```
5. Exit from the database prompt with `Ctrl + D`
6. Try re-running the problematic jobs

#### Option 7: Fully reset the SSD cache system

Fully reset the caching system with the following steps:

1. Ensure no jobs are running in CryoSPARC
2. For each connected worker machine:
   * Navigate to the SSD cache directory containing CryoSPARC's cache files (e.g., `/scratch/`). This path was configured during installation time
   * Look for a directory named `instance_<master hostname>:<master port + 1>` e.g., `instance_localhost:61001`
   * Delete this directory and all its contents

*The following additional reset instructions are required for CryoSPARC v4.4 or older.*

1. In a terminal, enter `cryosparcm mongo` to enter the interactive database prompt
2. Enter the following command to clear out the cache records

   <pre class="language-javascript"><code class="lang-javascript"><strong>db.cache_files.deleteMany({})
   </strong></code></pre>
3. Exit from the database prompt with `Ctrl + D`
4. Try re-running the problematic jobs

#### Option 8: Disable SSD cache for the job

If the issue persists after trying any of the above options, consider disabling the "Cache particle images on SSD" parameter for the affected job.

## User Interface Error Logging

If you encounter a problem with the user interface in your web browser, e.g., one or more elements of a page are not loading, etc., you can use the following steps to obtain debugging information.

1. **Open the browser console**
   * In Chrome, Firefox, Edge, and Safari this can be done by right clicking the page to open the browser context menu and selecting the `Inspect` option (`Inspect Element` in Safari) .
   * This will open up a “DevTools” panel used for inspecting and debugging in the browser. This panel includes a number of tabs at the top used to display different views. When opened using the context menu the current view will be the `Elements` tab. Click on the `Console` tab directly beside the `Elements` tab in order to view the web console. This is where errors, warnings, and general information about the page’s javascript code can be observed.
   * In order to keep the console clean in production we disable our development logs. Enable these logs by pasting the command

     `window.localStorage.setItem('cryosparc_debug', true);` into the browser console and then pressing the enter key on your keyboard. `undefined` will be logged below this command if it was submitted correctly.
   * Now reload the page and all of the development console logs and errors will be visible.
2. **Save Console Output**

   Please save console output as a `.log` file, including the type of browser (Chrome, Firefox, Edge, Safari, etc.) in the file name. The filename should be formatted as such: `console_{browser}.log` , eg. `console_chrome.log` . Before saving the file, try to reproduce the issue you encountered.

   * *Chrome or Edge:* Right click anywhere in the console panel to open the context menu and select the `Save As...` option. This will allow you to save the entire output as a `.log` file.
   * *Firefox*: Right click on a console message to open the context menu and select the `Save all Messages to File` option. This will allow you to save the entire output as a `.txt` file as default (or `.log` file optionally).
   * *Safari:* Click and drag the cursor over all items in the console output so that the items are all highlighted blue. You can then right click on any of the highlighted items to open the context menu and select the `Save Selected` option to to save the entire output as a `.txt` file as default (or `.log` file optionally).
3. **Save Network Output**

   Navigate to the `Network` tab in the DevTools by selecting it from the tabs in the top bar of the panel. If the `Network` tab is not shown then it is likely hidden in the overflow menu (this appears when there is not enough space to display all of the tab options in the DevTools). Click the overflow menu button represented by two right chevrons (`>>`) and select the `Network` option.

{% hint style="info" %}
Before saving the file, make sure to reproduce the issue that you encountered.
{% endhint %}

*Please include the type of web browser (Chrome, Firefox, Edge, or Safari) in the name of the* `.har` *file you are saving. The filename should be formatted as such: `network_{browser}.har` , eg. `network_chrome.har` .*

* *Chrome or Edge:* Click the "Export HAR" button on the panel header.

  <figure><img src="/files/SHpuuswpvcI8KWLEIXGb" alt=""><figcaption></figcaption></figure>
* *Firefox:* Right click on any of the items in the network request table and then select `Save All As HAR` from the context menu.
* *Safari:* Right click on any of the items in the network request table and then select `Export HAR` from the context menu.

## Additional Help

For topics not covered above, get additional help through the CryoSPARC Discussion Forum:

{% embed url="<https://discuss.cryosparc.com/>" %}

If no related discussions exist, please create a new post. Review our Troubleshooting Guidelines for items to include in your post:

{% embed url="<https://discuss.cryosparc.com/t/before-you-post-troubleshooting-guidelines/2355>" %}


# A Tour of the CryoSPARC Interface

{% hint style="warning" %}
The information in this section applies to CryoSPARC v4.0+.\
\
For the Application Guide applicable to CryoSPARC ≤v3.3, please see: [v3 User Interface Guide](/guides-for-v3/user-interface-and-usage-guide)
{% endhint %}

## Overview

{% embed url="<https://youtu.be/b9UD7_GGZas>" %}
CryoSPARC v4.0 Application Walkthrough.
{% endembed %}

The CryoSPARC application is a web interface that makes it easy to quickly and efficiently process cryo-EM data from raw movies to a high resolution structure. Within the interface you can create projects and workspaces to organize data, queue and run jobs, view and share their results and outputs, and export information for record keeping and use in other software. CryoSPARC also features additional tools to organize, search, and view data in a variety of different ways. This section of the guide will provide an overview of all these features.

<figure><img src="/files/PJfHID5UolU7RH4LheFJ" alt=""><figcaption></figcaption></figure>

## Application Layout

<figure><img src="/files/3sh1dSuwigkq6WUJXyl7" alt=""><figcaption></figcaption></figure>

The interface is comprised of five primary elements:

1. The navigation bar is located on the left side of the screen. It contains links to the homepage (Dashboard), the project browse view, and the session browse view. It also includes buttons to open various management dialogs, such as the 'current jobs' dialog and spotlight search dialog.
2. At the top of the page are the navigation controls, used to navigate to and switch between different view levels (projects, workspace, sessions, jobs). It is similar to the navigation controls in CryoSPARC v3 with added clarity and utility.
3. Below the navigation is a large content area that will adapt based on the page you're viewing. Every page has a centralized action bar with key elements such as filter controls.
4. Below the content area is a footer designed to provide quick access to current (queued or active) jobs. When browsing projects, workspaces, sessions, or jobs, the footer adapts to display a total count and various filter options.
5. To the right of the content area is a sidebar with three tabs: details, builder (for building and editing jobs) and cart (for filtering and creating jobs based on the outputs of completed jobs). Similar to CryoSPARC v3, you can select cards to view their details and perform actions. Additionally, you can now collapse the sidebar in order to view more of the main content area.

## Navigation Bar

<figure><img src="/files/luaDNOK5SjZR6zRDCqrN" alt=""><figcaption></figcaption></figure>

The navigation bar is located on the left side of the screen. It contains links to the:

* Dashboard
* Projects browser (all projects)
* Session browser (all sessions)
* Management dialog containing several tabs for managing jobs, data and instance information
* An overflow menu with additional dialogs and actions

## Dashboard

<figure><img src="/files/ZgdJmxodIfIdSNLvN4cX" alt=""><figcaption></figcaption></figure>

The first page you'll see after logging in is the Dashboard. It displays an overview of your CryoSPARC instance with helpful modules such as a processing history heat-map and charts to get a sense of what has been run recently as well as a module to see all active jobs. You can filter these modules to see information about only your jobs or jobs across all users within the instance. The Dashboard also includes additional modules featuring external CryoSPARC and community resources such as links to tutorials on the CryoSPARC Guide, trending posts on the CryoSPARC Discussion Forum and the latest EMPIAR and EMDB uploads.


# Browsing the CryoSPARC Instance

In order to interact with CryoSPARC and process data, you will need to browse the instance, navigating through the various projects, workspaces, sessions, and jobs that you and other users in the instance create. The Browse System is a set of tools and interfaces that help move around the instance and find the items you are looking for.

<figure><img src="/files/KelL9ab8X1iYn4wrGaYJ" alt=""><figcaption></figcaption></figure>

The browse system contains five fundamental components:

1. **Browse Header:** The header sits above the main working area and provides options for high level navigation and creation.
2. **Control Bar:** The control bar is where all actions involved in filtering, sorting, or setting a view option reside.
3. **Content Area:** The main space in the browse system is the content area. This is where projects, workspaces, sessions, and jobs are shown in either card, table, or tree view.
4. **Footer:** The footer shows total counts, curated filter options, and can be toggled using the footer switch button on the far left of the bar to show general active job and target information.
5. **Sidebar:** The sidebar shows contextual information and applicable actions for any selected items.

From the CryoSPARC dashboard, browsing the instance begins with clicking on the container icon in the left-side navigation bar. This button brings you to the project cards view, showing all projects you have access to across the instance.

## Navigation

### Item Cards

<figure><img src="/files/XIzHB6uSyHkiy6kzyCSZ" alt=""><figcaption></figcaption></figure>

When browsing through projects, workspaces, sessions, and jobs, each item will be displayed as a card by default.

Projects, workspaces, and sessions have buttons at the bottom of the card which can be clicked to open that item and navigate into a view of its contents. Jobs cards can be opened by clicking on the actionable header button which will expand the job inspection dialog.

Each type of card can also be selected by clicking on it, which will highlight it in blue and populate the sidebar with its information. From here the item can also be entered by clicking the “View” button at the bottom of the sidebar, or by pressing the `enter` key on your keyboard for projects, workspaces, and sessions, or the `spacebar` for jobs.

### Quick Switchers

<figure><img src="/files/m2e3g4gaCru5Ubu3DMdc" alt=""><figcaption></figcaption></figure>

The quick switchers at the top left of the page indicate your current location within the instance, and allow for quickly entering, exiting, and switching between projects, workspaces, and sessions. The solid blue switcher indicates the type of items you are currently looking at in the main content area.

These switchers can be used to quickly jump around the instance by clicking the arrow button on the righthand side of the switcher to open the selection dropdown.

Once you navigate inside a project, workspace or session, an “X” button will appear on the lefthand side of the corresponding switcher. Clicking the “X” button will deselect the option and navigate you back up one level of navigation, while clicking the name will reopen the selection menu and allow the quick selection of a different option.

### Quick Access Menu

<figure><img src="/files/3buey6jFkXQ1cNWWA8wB" alt=""><figcaption></figcaption></figure>

The Quick Access Menu is available by clicking on the menu icon in the bottom left corner of the page.

The menu is composed of three tabs: Recent, Starred, and Tags. Use these menus to quickly jump to important items in your instance.

* **Recent:** recently accessed items. Clicking on a project or workspace will open it and navigate you into the relevant section of the browse system. Clicking a session will navigate you into that session in the CryoSPARC Live view. Clicking a job will open its dialog without navigating away from your current location.
* **Starred:** items that you have starred. Starring items can be done directly from the items card’s header, and for jobs can be done in the header section of their dialog as well.
* **Tags:** all of the available tags across the instance along with counts of how many items they have been applied to. Clicking items will send you to a browse view with the selected tag applied as a filter. This quick access menu tab is covered in more depth in the “Tags” section below.

{% hint style="info" %}
The last menu tab you accessed is remembered by the browser in the current browser tab until closed. This is in order to aid in navigating swiftly between items even when closing and reopening the menu. Each sub section inside of each menu tab is a drawer that can be closed by clicking the section header (e.g., Recent Projects). When an item is closed it will remain closed between browser tabs and even if the browser is closed and reopened. This allows you to organize each tab of the quick access menu to only show the items that are most important to you.
{% endhint %}

### Spotlight

<figure><img src="/files/JWIF0lbt8ihPKwHZIjHo" alt=""><figcaption></figcaption></figure>

The spotlight allows unrestricted navigation through the interface and can be opened by clicking the magnifying glass icon in the navigation bar, the search bar button on the home page header, or by pressing the `command` + `k` keys together on the keyboard.

The spotlight input is automatically focused and you can begin typing into it immediately. Many commands are available, such as creating a new item, opening the jobs builder, or attaching a project.

Commands specifically for navigation include viewing all items in a granularity, or typing in the name of a project, workspace, or session to quickly find and open them.

## Views

When browsing the CryoSPARC instance, you can select a view from the buttons in the top right corner of the content area.

### Cards View

<figure><img src="/files/tujI10E01xmZptWbmWgG" alt=""><figcaption></figcaption></figure>

Items in CryoSPARC display as cards by default. Each card includes a header with the ID and title of the item along with action buttons for opening the card’s action menu and starring the item.

The card can be selected by clicking on it anywhere that is not an actionable button. This will add a blue selection outline to the card and add all of its relevant information to the sidebar. Once selected, the item can be opened by pressing the `enter` button on your keyboard (for projects, workspaces, and sessions). Items can also be opened by clicking the “view” button(s) at the bottom of their respective card, or by clicking the view button at the bottom of the active sidebar. Entering the selected item will navigate into the item and show all of its contained items (eg. entering a project will navigate to the workspaces view and show all of the workspaces belonging to that project). For jobs, selected cards can be opened by pressing the `spacebar` to expand the corresponding job dialog. Job cards can also be opened by clicking the card header button with the job ID and name.

### Table View

<figure><img src="/files/J3N3mRPYMkwP4MJCydFT" alt=""><figcaption></figcaption></figure>

The table view organizes each item as a row in a table. This can be particularly useful for getting an overview of large numbers of items and for accounting purposes. When the row is selected it will highlight blue and can be opened the same way as with a card, by pressing the enter button on your keyboard or by clicking the view button at the bottom of the sidebar. The download button at the bottom right of the jobs footer allows for exporting the entire data selection with the same filter, sorting, and granularity options as the table view in a CSV file.

### Tree View

<figure><img src="/files/Q5H3xCLoPE0h9W30tqwf" alt=""><figcaption></figcaption></figure>

The tree view is a unique view only available for jobs inside of a workspace or project. This view lays out all jobs as an interconnected network branching out from the first job(s) in the processing pipeline. Each job is linked to previous and subsequent jobs by coloured lines representing the flow of inputs and outputs between jobs.

Navigate the tree view using the two control palettes present in the bottom left and right corners of the viewing area. Zooming is the default scroll-wheel behaviour, which can be changed by clicking the “pan” icon in the bottom right corner of the tree view. Default zoom levels are shown in the bottom left corner of the viewing area. These can be used to quickly toggle between different levels of zoom to quickly move the entire view in or out to a set scale factor. When first switching to the tree view, all trees will be centred within the view both vertically and horizontally, making it easy to locate your data immediately upon entry.

Tree view job cards contain a more condensed version of the job data available in the standard cards view. Available actions are shown on hover in a command palette located in the bottom righthand corner of each card. Right clicking the card will show the quick actions menu just like cards from the card view and rows in the table view. Additional information is shown in a tooltip that will appear above the card when it is hovered for a short period of time. These tooltips along with the quick actions menu make it easier to see relevant job details and act on the job without needing to change your zoom level in order to do so.

{% hint style="warning" %}
Due to the unique structure of the tree view, sorting actions cannot be applied in this view.
{% endhint %}


# Projects, Workspaces and Live Sessions

{% hint style="danger" %}
Do not remove from the filesystem any directory that is managed by an[ attached CryoSPARC project](https://guide.cryosparc.com/application-guide/pages/F3KBgDxkuaoVRFwpV0KW#2.-attaching-detaching-archiving-and-unarchiving-projects). First delete unwanted projects using the *Delete Project* GUI action or the [`delete_project()` method of the CryoSPARC CLI](/setup-configuration-and-management/management-and-monitoring-4.7/cli-4.7#delete_project-project_uid-str-request_user_id-str-all_jobs_in_project-list-all_workspaces_in_projec).
{% endhint %}

Projects in CryoSPARC are high level containers corresponding to a project directory on the filesystem which houses all associated jobs. Each project in CryoSPARC is entirely contained within the project directory. All of the jobs and their respective intermediate and output data created within a Project will be stored within the project directory.

Projects are strict divisions. Files and jobs from different projects are stored in dedicated project directories and jobs cannot be connected from one project to another.

Workspaces, on the other hand, are logical groupings like labels, that are created by the user to separate portions of a workflow for ease of use. A job can be added to multiple workspaces at the same time and can be removed from a workspace at any time. Workspaces do not have any particular directory on the filesystem and there are no file transfers or copies made when jobs are moved between workspaces.

## Projects

### When to create a new project

{% hint style="info" %}
Recommendation: Create a new project for each new unrelated sample on which you are collecting data.
{% endhint %}

Additional recommendations for creating new Projects and Workspaces:

* **Collecting new data for the first time on a new target molecule:** Create a new project and a new workspace within it. Import the movies/micrographs/particle stacks into the new workspace.
* **Collecting data a second or subsequent time on the same sample/target** (potentially the same or different grid from the same batch, potentially on a different day): Use the existing project and existing workspace where you processed the first set of images. Import the movies/micrographs/particle stack into the existing workspace or a new workspace, in the same project.
* **Collecting data on a new sample/grid/preparation of the same target molecule:** Use the existing project, but create a new workspace. This allows easy re-use of 3D volumes, 2D templates, and easy combining of particle images downstream. You can create multiple workspaces within a project, for example if collecting/processing new data from a similar sample.

See similar considerations for creating projects and sessions for CryoSPARC Live in the [New Live Session: Start to Finish Guide](/live/new-live-session-start-to-finish-guide).

### **Creating your first project**

1. Navigate to the Projects view by clicking on the container icon on the left-side navigation bar
2. To create a project, click the green “New Project” button in the top right of the browse header (above the main content area). This will open the “slide-over”panel where you can enter you project details and create your new project. Alternatively, if no projects currently exist, the “Create a New Project” panel will appear in the main content area. This functions identically to the slide-over panel.
3. Enter a project title and select a container directory, which is the location that a new project directory for the new project will be created. The container directory you select should already exist. The new project directory will be created and will be populated with job directories as you create jobs. All files associated with the project will be stored inside the project directory. You may also wish to enter a description for your project.
4. Clicking the “Create” button at the bottom of the slide-over will create your new project and add it to the project page. This will also automatically take you into the created project where you can then create your first workspace.

<figure><img src="/files/vd7MaUSXATdfy4i7VPr9" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/ZUEdTPJqkney7tRyXGQ4" alt=""><figcaption></figcaption></figure>

A project can be opened either by clicking the navigation button at the bottom of a project card, by selecting the project and pressing the enter key, or by clicking the “View Project” button at the bottom of the sidebar. Alternatively, a project can be selected using the quick switcher at the top of the page. Any of these options will open the corresponding project and show all of its workspaces.

## Workspaces

### Creating your first workspace

Before you can start processing data inside of a new project, you will need to create at least one workspace.

1. If you have not already opened your project, you can do so from the projects view by clicking on a project to select it and either pressing the `enter` key, or pressing the “View Project” button at the bottom of the active sidebar. Alternatively, you may simply click the “View x Workspaces” or “No Workspaces” button (which will show when no workspaces exist inside of your project yet) at the bottom of the project card.
2. Once inside your selected Project, create a New Workspace by clicking on the green “New Workspace” button in the top right of the browse header, or simply entering the relevant information into the “Create a New Workspace” panel that shows inside of an empty project. Alternatively, you can open the actions panel by clicking the “Actions” button at the bottom of the project sidebar and then clicking the “New Workspace” button. Any of these actions will open the workspace creation slide-over panel. Workspace titles can be changed later, and descriptions can be added any time.
3. Clicking the “Create” button at the bottom of the slide-over will create your new workspace and add it to the workspaces view. This will also automatically take you into the created workspace where you can begin creating jobs and start processing data.

<figure><img src="/files/2SI5KzHP86oGof75tWw9" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/f8DR82w358qvYwIqNNLc" alt=""><figcaption></figcaption></figure>

## Project and Workspace Interface Details

### Project Card

<figure><img src="/files/PUSC0QfKMM3eVOoAaFUL" alt=""><figcaption></figcaption></figure>

The project card is composed of four main sections:

**Header**

The header shows the project’s ID and title as well as a context menu trigger and a star button (visible when the card is hovered).

**Preview Panel**

The job preview track shows jobs within the project sorted by most recent. These preview cards can be clicked to open the jobs that they represent. They also allow for a quick way to see what was last being worked on within the project.

**Information Panel**

This panel displays a variety of distinct information that can be useful at a glance.

* The date when the project was created and the user who created it.
* A list of tags associated with the project. Tags can be created and added to all browse items and can be used to organize and filter data in the browse system.
* The description added to the project if applicable. Descriptions are markdown compatible.

**Navigation Buttons (Bottom)**

This is the the primary way of entering the project, and also displays the number of workspaces/sessions existing within a project.

### Project Row

<figure><img src="/files/XIFBa9vWkNHNoMqiGZo5" alt=""><figcaption></figcaption></figure>

The project row shows in the Table view, and gives a less information dense overview of the project allowing for easier navigation and management of projects on larger instances. It includes counts of all workspaces, sessions, and jobs inside of the projects, as well as a description cell with a markdown formatted tooltip.

### Project Sidebar

The right-side sidebar, when a project is selected, shows details about the project. It is composed of a few sections:

#### **Header**

The sidebar header shows the project ID, project title, and includes a deselect button on the righthand side. This button will remove the current sidebar selection and fall back to displaying details of the container you are browsing (eg. if a workspace is selected and you are inside of a project, deselecting the workspace sidebar will replace it with the details of the project).

<figure><img src="/files/qULnCsv2paDyRNSIUqSC" alt=""><figcaption></figcaption></figure>

#### **Details**

<figure><img src="/files/RygE3BTPePXlpHbPydUZ" alt=""><figcaption></figcaption></figure>

The details panel shows a variety of core information and statistics such as:

* **Title**: The unique title given to the project to identify it.
* **Tags**: All tags that have been applied to the project. This section has an edit button visible when hovering over it. Clicking this button will open up a context menu where you can select more tags to add to the project, or deselect added tags to remove them. Clicking the coloured tag pill buttons will add that tag as a filter to the current view.
* **Created:** This is the timestamp indicating the date and time when the project was initially created.
* **Created By:** The name of the user who initially created the project.
* **Last Accessed:** The timestamp indicating the last time the project was opened. This time can represent access by any user who is able to view the project.
* **Last Accessed By:** The full name of the last user who opened the project.
* **Directory:** The absolute path to the filesystem directory where the project is located. This path can be copied by clicking on it.
* **Size:** The total size of the project on disk in gigabytes.
* **Size Last Updated:** This represents the last time the project size was calculated. CryoSPARC regularly refreshes statistics across the instance on a set schedule to keep them up to date. This limits heavy calculations that would need to be made every time a statistic is viewed. The size displayed can be manually refreshed by clicking the “refresh” button beside the timestamp. This will force the system to recalculate the project size to show the absolute most up to date version.

#### **Description**

<figure><img src="/files/DosEtmPEURHmTPs7pWDz" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/yIVUs5yiFzo2bEO9oYP6" alt=""><figcaption></figcaption></figure>

If no description is present this panel will default to edit mode and display a small markdown editor that saves text automatically as you type into it. If a description has been added to the project this panel will show a markdown formatted read-only pane with an “Edit” button in the top right corner visible on hover.

#### **Sharing**

<figure><img src="/files/ntOaQWC6z0uODO0Cp2Q5" alt=""><figcaption></figcaption></figure>

The sharing panel allows you to easily see who you have shared your project with, as well as share with new users and remove users who you no longer wish to share the project with. Clicking the share button at the bottom of the panel will open a context menu populated with the names of other users across your current instance. Clicking any of these users will share the project with them, allowing them to see and modify data inside of it. The selected users will show up in the panel as new rows below the owner row (your user). Each of these rows has a remove button on the righthand side that can be clicked to remove sharing privileges from the user. Removing a user will stop them from being able to see or interact with your project in any way.

#### **Statistics**

<figure><img src="/files/iwFLlj32hZGzmHSs8OGR" alt=""><figcaption></figcaption></figure>

* **Last Updated:** Shows when the projects statistics were last updated and includes a button beside the timestamp allowing you to manually refresh them if they are stale.
* **Workspaces:** Shows the number of workspaces in the project, and allows quick navigation into the project to the workspaces view by clicking the “view” button visible on hover.
* **Sessions:** Shows the number of sessions in the project, and allows quick navigation into the project to the sessions view by clicking the “view” button.
* **Jobs:** Shows the total number of jobs in the project across all workspaces, and allows quick navigation into the project to a flattened jobs view by clicking the “view” button. This view shows all jobs across the project regardless of workspace.
* **Job Status:** Each coloured status button in this section shows the total number of jobs inside of the project with the corresponding status. Clicking any of these buttons will navigate into the project to the jobs view and show all jobs inside of the project filtered by the selected status.
* **Job Types:** This section shows all of the different job types across the project and a count of the number of jobs corresponding to each type. Each job type indicator is a button that can be clicked to navigate into the project jobs view with a filter of that job type applied (eg. only jobs of that type, from across the entire project, will show in this view).

**Project-Level Parameters**

<figure><img src="/files/IIxaobet7ofUCYqPqkQx" alt=""><figcaption></figcaption></figure>

Enable SSD Caching: This section includes a selection menu with options for setting a project level override for the top level SSD Caching parameter. This selection will only take effect on the specific project it is set on.

#### **Actions**

<figure><img src="/files/8fVYKdtdH4vy92iCdPqT" alt=""><figcaption></figcaption></figure>

The “Actions” button at the bottom of the project sidebar will open a slide up command palette when clicked. This palette contains a number of relevant actions that can be taken on the project.

* **New Workspace:** This will open the slide-over panel for workspace creation. Filling in the required fields and clicking the create button will make a new workspace inside of the project and navigate you to the project’s workspaces view with the new workspace selected.
* **New Session:** The same functionality as the new workspace action except for creating a live session.
* **Archive Project:** Allows the project to be archived to remote/slow storage to free up space. A record of this project will remain in your projects view with an “archived” tag.
* **Clear Intermediate Results:** This will clear all unused intermediate results from all jobs within the project to free up disk space.
* **Detach Project:** Removes the project and its data for use in attaching to another instance. Detached projects are displayed in the application for reference only. No actions can be taken on a detached project.
* **Delete Project:** This action removes all data generated in that project, including micrographs, extracted particles, and any other associated data. Deleting a project does not delete the raw data that was imported into the project.

### Project Quick Actions Menu

<figure><img src="/files/V4AMaW6UzSAjJpAAGAqC" alt=""><figcaption></figcaption></figure>

The quick actions menu is accessible on a project by right clicking on a project card or row, or by clicking the button with three horizontal dots on the card header. This will open a context menu with a variety of actions that can be taken on the project. These actions are largely identical to those present in the “Actions Palette” accessible from the sidebar, a reference of which can be seen directly above this section.

Along with the aforementioned actions, the quick access menu also allow tags to be quickly added or removed from the project using a submenu accessible from the “Edit Tags” menu item.

### Workspace Card

<figure><img src="/files/JlwY2HtTt0P8EP7x6wS0" alt=""><figcaption></figcaption></figure>

The workspace card is largely analogous to the project card in terms of composition of available information and actions.

One additional section is available in the workspace card’s information section, an actionable list of job categories with counts of each job within the workspace that fall into said category. Each item in this list is a button that when clicked will automatically open the workspace and show all of its jobs with the relevant filter applied (eg. clicking on a “Particle Picking” category count button will open the workspace and show only the jobs that fall within that category inside: ex. Blob Picker, Inspect Picks, Extract From Micrographs, Template Picker).

Opening the workspace, either by clicking the “view” button or selecting the workspace and pressing the enter key, will take you to the jobs view and show all jobs within.

### **Workspace Row and Sidebar**

The workspace row and sidebar are almost entirely identical to those corresponding to the project, but with a subset of data relevant to the workspace. Refer to the project section above for more information regarding these interface components.

### **Workspace Quick Actions Menu**

The same as with projects, the quick actions menu is accessible on a workspace by right clicking on a workspace card or row, or by clicking the button with three horizontal dots on the card header. This menu shows the same actions as in the sidebar’s “Actions Palette”, along with a submenu for adding and removing tags from the workspace.

## CryoSPARC Live Sessions

A CryoSPARC Live Session is actually a special type of workspace. It contains extra metadata and is used primarily for processing data within CryoSPARC live. Sessions have all the same functionality as standard workspaces and therefore are also visible in the workspace section of the browse system.

### Sessions view

The sessions view is a special view that shows Live sessions only, and displays them in a way that is more helpful when processing data in Live. The sessions view can be accessed by clicking on the lightning icon on the left-side navigation bar.

This view differs slightly from that of the other browse pages in that it includes a distinct session projects sidebar on the lefthand side of the main content area, and larger item cards to accommodate the extra data useful for identifying and analyzing live sessions at a glance.

<figure><img src="/files/Pc5EV6aXOtaEv5Rq2kXm" alt=""><figcaption></figcaption></figure>

### **Session Projects Bar**

The projects bar (left side of the content area when in sessions view) shows all projects that include live sessions for easy navigation. Each project card within the bar displays that project’s ID, name, the user that created it, how many sessions exist inside of it, and the date it was created. These project cards can be selected to filter sessions belonging only to the selected project or projects. Clicking a selected project card will deselect it, and if it was the only selected project then all projects will be selected instead.

### **Session Card**

<figure><img src="/files/Gz6wRSfXDOKxFxS50Qnw" alt=""><figcaption></figcaption></figure>

**Header**

The session card header is nearly identical to the workspace card header, with the difference of showing the distinct session ID and the current processing status (running, paused, or completed).

**Preview Panel**

The lefthand side preview image shows an exposure preview that can be clicked on to zoom.

The righthand side shows the job preview track, displaying jobs in the session sorted by most recent (the same view as in the project and workspace cards).

**Session Details Panel**

* **Exposures:** Shows the total overall exposures, completed exposures, queued exposures, and failed exposures.
* **Particles:** Shows the total manual picks, blob picks, template picks, and extracted particles.
* **Reconstruction:** Shows the current reconstruction resolution and all jobs run in the live reconstruction pipeline. The displayed jobs are buttons that can be clicked on to view the full job dialog.

**Details Panel**

The details panel shows the same information as the workspace card - creation information, description, job categories, and tags.

**Action Buttons**

* **View Session**: This button will open the session in the CryoSPARC Live interface
* **View Jobs**: This will enter the session and show all the applicable jobs, the same as entering a workspace.

Sessions also support entry by using the same keyboard shortcuts as project and workspace selections (pressing `enter`). Similarly, the session has its own action menu accessible from the header or by right clicking on the card or row.

Viewing the session will open the CryoSPARC Live view for that session. Viewing the session’s jobs will open the session and show all jobs associated with it.

For comprehensive documentation on CryoSPARC Live, please see the guides at:

{% content-ref url="/pages/-MNinyJcDXq4SuvhpE7W" %}
[About CryoSPARC Live](/live/about-cryosparc-live)
{% endcontent-ref %}

## Project and Workspace Management

### Share a Project

Only users who own a particular Project or have a Project shared with them can see Workspaces, CryoSPARC Live Sessions and Jobs within those Projects. To add a user to a Project, navigate to the Project Details Panel for a particular project (owner should be the logged in user) and click `Share With Users` to select the user you wish to give access.

### Delete a Project

On the Project Details Tab, the Delete Project button allows deleting an entire project. A pop-up message will ask you to confirm the action. This action removes all data generated in the project, including preprocessed micrographs, extracted particles, output volumes, and any other associated data. Deleting a project does not delete the raw data that was imported into the project. Note that deleting a project deletes the contents of the project directory but does not delete the project directory itself.

### Delete a Workspace

Navigate to the Workspace and locate the 'Delete' button on the Workspace Details Tab. A pop-up message will ask you to confirm the delete. Once you confirm, another popup will show you two lists: A. jobs that reside only inside the selected workspace, and that will be fully deleted, and B. jobs that exist in the selected workspaces, and other workspaces (i.e., because they were [linked](https://guide.cryosparc.com/guides-for-v3/user-interface-and-usage-guide/queue-job-inspect-job-and-other-job-actions)). Jobs that exist in multiple workspaces (list B) will not be deleted. They will simply be removed from the workspace to be deleted, and will still be accessible in the other workspace(s) where they were linked. For the jobs that are deleted, all of their intermediate and final output data is deleted from the filesystem.<br>


# Jobs

The fundamental unit of data processing within CryoSPARC is a job. Jobs are given inputs and parameters, do work, and through that work, create outputs. The outputs of one job can be connected as the inputs of another job, and through this process, allow you to create branching pipelines through which you can explore and refine your data.

## Job Card

The most common representation of a job in CryoSPARC is a “Job Card”. This card is composed of a number of different sections used to organize job information in a consistent way with flexibility to accommodate the specific differences between job types.

<figure><img src="/files/2ekrPDPwfGb3jnou6ctX" alt=""><figcaption><p>A 2D Classification job in the default image view</p></figcaption></figure>

### Header

The header shows the job’s ID and a status indicator with the colour of the job’s current status (e.g. green for completed, red for failed, etc). The whole lefthand side of the header is a button that will open the job dialog when clicked.

The righthand side of the header shows all of the available job actions, including buttons to toggle the outputs view, show the quick actions menu, and a button to star the job for future reference.

### Central Area

Underneath the header, you can give your job a unique title if you wish.

The description is an optional text field that can be added and edited in the sidebar (pencil icon). This field supports markdown and will turn into a scrollable sub-section on the card if a non-trivial amount of information is added.

The corners of the job card are populated with at-a-glance information - [info tags](https://guide.cryosparc.com/application-guide-v4.0+/inspecting-data#job-info-tags) - about the outputs, such as the number of particles or classes. This view allows you to get a quick overview of your job without needing to open it.

### Outputs View

The outputs view can be accessed either by clicking the outputs view toggle (cube with arrow icon) in the card header, or by pressing the `O` key to switch to the view on all cards at once. This view shows a list of all of the outputs created by the job. These outputs can be used in two ways:

<figure><img src="/files/7NWhcaBlLHLyJdFYQjwA" alt=""><figcaption><p>A 2D Classification job in the “outputs” view. Outputs are selectable for use in the "Job Builder" and “Job Cart”.</p></figcaption></figure>

* Outputs can be dragged and dropped into the inputs of a building job, visible in the sidebar when the job is selected. Further information is available in the “[Tutorial: Job Builder](https://guide.cryosparc.com/processing-data/tutorials-and-case-studies/job-builder-tutorial)” guide page.
* Outputs can alternatively be clicked, which will add them to the “Job Cart”, a system that allows you to filter sequential jobs if their inputs match the selected outputs. This system can be further explored in the “[Creating and Running Jobs](https://guide.cryosparc.com/application-guide-v4.0+/creating-and-running-jobs)” section of the guide.

### Footer

The job’s footer contains a variety of information to gain a better idea of the processing that has occurred within a job at a glance. Each icon can be hovered to reveal a tooltip with additional relevant information.

* The timer shows how long your job has been running or how long it ran for. It can be hovered to show who created the job and when they created it, along with the time it concluded processing if it is completed, killed, or failed.
* The lane icon shows the name of the lane to which the job has been assigned. Hovering will show the target name and type.
* The GPU icon shows the number of GPUs used by the job, and if running, will show the names of the GPUs in use. Hovering will show the number of GPUs used and available on the target.
* The gear icon shows the number of custom parameters set on the job, hovering shows those parameters and their values.
* Tag buttons shown beside the tag icon represent the tags that have been applied to the job. Clicking on these tag buttons will set the given tag as a filter in the current view.

## Job Actions

There are a variety of actions available for interacting with jobs. These range from directly assigning jobs the computational resources that they need to run, to organizing jobs between different workspaces, and even building templates from pre-existing jobs to expedite your work.

Job actions can be accessed in two places: the “Quick Actions Menu”, and the sidebar “Actions Menu”.

### The “Quick Actions Menu”

You can access the quick actions menu by clicking the trigger button in a card header (the trigger button is represented by an icon with three horizontal dots) or by right clicking on the card.

<figure><img src="/files/wJn1SdY2kObW7S8F3eRG" alt=""><figcaption></figcaption></figure>

### The sidebar “Actions Menu”

This menu can be accessed by clicking the “Actions” button at the bottom of the job sidebar with the job you wish to perform actions on selected. This menu contains all job core and compound actions, but does not include quick actions, selection actions, or navigation actions.

<figure><img src="/files/otvGvQJFlfWQyMIdAj0a" alt=""><figcaption></figcaption></figure>

### Core Actions

* **Queue Job**: Queuing a job will assign the necessary computational resources to it so that is can begin processing your data.
* **Link Job**: Linking a job allows you to connect a job in one workspace to another workspace within the same project. The job will exist in both workspaces as though it were created in them.
* **Move Job**: Moving a job will connect it to another workspace, and remove it from the current one.
* **Clone Job**: Cloning a job will create an identical copy of it, this includes all of its custom parameters and inputs.
* **Kill Job**: Killing a job will interrupt its processing and completely stop it from running. The job will release its computational resources and tokens.
* **Clear Job**: Clearing a job resets it to building status so that its parameters and inputs can be edited again. This is necessary to return the job to an operable state after it has failed or has been killed.
* [**Clear Intermediate Results**](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/tutorial-data-management-in-cryosparc#id-4.-ability-to-clear-intermediate-results): This will clear all unused outputs created by the job, in order to save storage space used by the raw data.
* [**Export Job**](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/tutorial-data-management-in-cryosparc#id-4.-ability-to-clear-intermediate-results): This will export all essential files (images, streamlog events, etc) into a folder inside of its parent project’s directory. This folder can be used to import the job into another project or even instance.
* **Mark Job as Complete**: Marking a job as complete will change its status to complete and allow you to access its outputs as though the job had completed normally. This can be useful if a job fails or is killed but the generation of its outputs had reached a point sufficient for further data processing.
* **Unlink Job from Workspace**: Unlinking a job will remove it from the current workspace. It will still be fully accessible in all other workspaces it had been linked to or from.
* **Delete Job**: Deleting a job removes it permanently from all views, and renders it inoperable for further data processing. Deleted jobs can still be accessed through a filter for accounting purposes, and continue exist in the database in a limited form.

### Multi Actions

*These actions can only be performed with multiple jobs selected.*

* **Clone {x} Jobs**
* **Clone Job Chain** (only available for valid chains of jobs): This action allows you to clone all jobs between, and including, two initial job selections. One job represents the start of the chain, and the other represents the end of the chain.
* **Delete {x} Jobs**
* **Create Workflow**: Workflows are a multi-job template that allow you to quickly build pre-populated sets of jobs to create automated pipelines. They can be explored in more detail in the “[Workflows](https://guide.cryosparc.com/application-guide-v4.0+/workflows)” section of the guide.

### Quick Create Actions

* Quick create actions are touched on in more detail in the “[Creating and running jobs](https://guide.cryosparc.com/application-guide-v4.0+/creating-and-running-jobs)” section of the documentation. These actions allow for the creation of predefined template jobs that are automatically connected to the outputs of the current job.

### Secondary Actions

* **Create Blueprint**: Blueprints are a per job template that can be used to quickly create a job and populate its parameters, or populate/overwrite the parameters of a pre-existing job. They are covered in more depth in the “[Blueprints](https://guide.cryosparc.com/application-guide-v4.0+/blueprints)” section of the guide.
* **Edit tags**: Allows you to add or remove pre-configured tags from the job. Tags and tagging are expanded upon in the “[Tags](https://guide.cryosparc.com/application-guide-v4.0+/tags)” section in the guide.

### Selection Actions

* **Select Job Chain**: Selects all jobs between two initial job selections, one represents the start of the chain, and the other represents the end of the chain.
* **Select Ancestor Jobs**: Selects the chain of jobs that lead up to and connect to the current job.
* **Select Descendent Jobs**: Selects the chain of jobs that connect to and branch out from the current job.

### Selection Actions

* **Select Job Chain**: Selects all jobs between two initial job selections, one represents the start of the chain, and the other represents the end of the chain.
* **Select Ancestor Jobs**: Selects the chain of jobs that lead up to and connect to the current job.
* **Select Descendent Jobs**: Selects the chain of jobs that connect to and branch out from the current job.

### Navigation Actions

* **View workspaces in {project}**: This action will navigate you to the workspaces page corresponding to the project that the job is inside.
* **View jobs in {project} {workspace}**: This action will navigate you to the jobs page within the workspace that the job is assigned to.


# Job Views: Cards, Tree, and Table

## Cards View

Job cards in CryoSPARC are displayed by default in a masonry grid referred to as the “Cards View”. Cards are laid out from left to right and then in sequential rows organized by a sort attribute and sort order. The default sort attribute is the “Date Created”, and the default sort order is ascending (this means the most recent jobs will appear at the bottom of the view, while the oldest will appear at the top).

<figure><img src="/files/6oR77Lrpd1fCP6Gmk6nw" alt=""><figcaption></figcaption></figure>

### Navigation

The cards view is a scrollable page that progressively loads job information as you navigate. This means that you can use the scroll bar or your mouse wheel to navigate to any point in the view, whether you are working on 100 jobs or 1000 jobs, without needing to change pages.

#### Targeting

The target button on the filter bar allows you to quickly navigate the view to the currently selected job. This can also be done by pressing the `T` key on your keyboard. The adjoined arrow button will open a dropdown menu with options for the “Last Running Job”, “Last Completed Job”, or “Last Created Job”. Selecting an option from this menu will select the corresponding job and then navigate you to it.

<figure><img src="/files/b6Nzfc86SIQVyZrW4vjt" alt=""><figcaption></figcaption></figure>

#### Searching

The job search menu can be opened by clicking on the job count button in the footer. This menu contains a list of jobs in the current view (workspace, project, or instance). The menu items include a colour indicator for the job’s current status, as well as its job ID and job type. The list can be filtered by typing into the input at the bottom of the menu. Clicking on a job in this list will select the job and then immediately navigate you to it.

<figure><img src="/files/EtiEtkOqMB8uAFCq5M2y" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
The job search menu is limited to the first 1000 jobs in the view, sorted by the attribute and order that you have selected.
{% endhint %}

### Filters

A variety of filters are available to help you quickly and easily find sets of jobs that match a specific criteria (eg. jobs that were created between two dates, or have a specific job type). All jobs that match the filter will continue to be displayed, and those that do not will be removed from the view. Filters can be applied additively to create a more granular criteria for matching jobs (eg. all Ab-Initio jobs that have a status of completed or failed). More information on the filter system can be found in the "[Filters and Sorting](https://guide.cryosparc.com/application-guide-v4.0+/filters-and-sorting)” section of the guide.

<figure><img src="/files/WtTSZ86Aa2HiTVB6aa9X" alt=""><figcaption></figcaption></figure>

## Tree View

The tree view is a unique view available for jobs inside of a workspace or project. This view lays out all jobs as an interconnected network branching out from the first job(s) in the processing pipeline. Each job is linked to previous and subsequent jobs with coloured lines representing the flow of inputs and outputs between them.

The tree view is in many ways a “superset” of the cards view, which is to say that the job cards themselves look and behave for all intents and purposes identically to the card view, and jobs also retain all of the same information and actions. All systems available in the card view for navigation (ie. targeting and searching) are also available in the tree view.

{% embed url="<https://youtu.be/HBG6t_aijo0>" %}

### Navigation and Selection

When initially navigating into the tree view, the view will load with all jobs visible and centred horizontally and vertically within the viewing area. The scroll wheel is set to zoom in and out of the view by default, and you can click and drag anywhere in the viewing area to navigate.

#### Zoom and Pan

The zoom or pan modes are shown in the bottom right of the view, as the magnifying glass and four-arrows icons respectively. These control the scroll behaviour. In zoom mode (default), scrolling will zoom the view in or out, while in pan mode, scrolling will pan the view horizontally or vertically. Holding the `command`/ `ctrl` key while scrolling will invert the action. For example, by holding down the command key while in zoom mode, scrolling would cause the view to pan instead. Both options are also mapped to their own keyboard shortcuts. Zoom mode is mapped to the `Z` key while pan mode is mapped to the `X` key beside it.

#### Drag and Select

The two options for the click mode, drag and select, are shown in the bottom right of the view as the hand and pointer icons respectively. By default, the click mode is set to drag. This mode allows you to click, hold, and drag the tree view to navigate around it. By switching to select mode you can click, hold, and drag the cursor to select multiple jobs in a rectangular selection box (much like drag selection of files on the desktop of most modern operating systems). Drag select will, by default, start a new selection each time you begin dragging the selection box. You can retain your current selection and select additional jobs by holding the `command` / `ctrl` key while making your selection. Drag mode is mapped to the `C` key while select mode is mapped to the `V` key.

#### Scale Levels

The scale switcher is located on the bottom left of the viewing area and allows you to reset the view to show all of the available jobs, select a custom zoom level, and cycle through multiple preset zoom levels. The `R` key will reset the view, while holding the `shift` key and then pressing the `R` key will cycle through the preset zoom levels.

### Filters

Filters can be applied in the tree view in the same way as in the cards view, but in order to retain the spatial context of job trees, filtered jobs will still appear in the tree view in the same position but as dotted outlines, with no input or output connections.

Jobs that are not filtered out appear as usual and will have their input/output connections rendered if they connect to other jobs that have also not been filtered out.

<figure><img src="/files/2WjiXwEi52zX1pKJbHKz" alt=""><figcaption></figcaption></figure>

### Upstream Jobs

To correctly render the tree view layout, it may be necessary to display job cards for jobs that are not part of the current workspace. These jobs are siblings of jobs within the workspace but have not themselves been linked to or moved into it.

Such jobs are referred to as **upstream jobs**. In the interface, upstream jobs are displayed either as grouped cards containing multiple jobs (the default presentation) or as individual job cards outlined in orange.

<figure><img src="/files/ST4n3EPKpr9hBxcaSyo1" alt=""><figcaption></figcaption></figure>

#### Grouped Cards

By default, upstream jobs are presented as grouped cards, each representing a variable number of jobs. These groups are generated automatically based on the proximity and connectivity of upstream jobs relative to the current workspace.

The card header displays a badge indicating the number of upstream jobs contained within the group and can be clicked to select all jobs in that group.

The card body consists of a scrollable list of the contained jobs. Each entry displays the job’s ID and type and can be selected individually. Jobs in this list that have a direct parent and/or child connection within the current workspace display an indicator badge on the right-hand side of the entry.

<figure><img src="/files/neuZyUrbn1SunanZDm30" alt=""><figcaption></figcaption></figure>

#### Single Cards

When expanded, upstream jobs are displayed as individual job cards. These cards are visually distinguished by a solid orange border and an unlinked icon in the footer, but are otherwise consistent with the appearance of standard job cards in the workspace.

Upstream job cards have several functional differences designed to reduce the risk of accidental selection or modification. They cannot be selected using the drag-selection tool and must be selected individually. Additionally, the quick actions menu is disabled as a context menu; actions must instead be performed by opening the menu via the card header or by using the sidebar actions panel.

<div data-full-width="false"><figure><img src="/files/wiNpRZ9P87b59ey3Iihx" alt=""><figcaption></figcaption></figure></div>

### Shortcuts

*Note: `command` key represents `command` on Mac and `ctrl` on Windows or Linux*

* `Z` key will set the scroll mode to zoom.
* `X` key will set the scroll mode to pan.
* `C` key will set the click mode to drag.
* `V` key will set the click mode to select.
* `T` will reset the view on the currently selected job or set of jobs.
* `R` will reset the view to show all jobs.
* `shift` + `R` will cycle through zoom levels (0.25x, 0.5x, 1x).
* Holding the `command` key while scrolling will invert scroll mode to pan or zoom (depending on which is currently selected).

## Table View

The table view presents jobs with their information laid out in a highly structured and predictable manner. The rows represent jobs, and the columns represent the different types of information that the jobs contain.

<figure><img src="/files/SqL9jpaKJWw2jAZ3oZmG" alt=""><figcaption></figcaption></figure>

The table view is particularly useful for accounting purposes, as it allows you to view a condensed representation of you job data, much like a spreadsheet.

All of the filters and sorting options that are available in the card and tree view are available in the table view as well. Jobs can be quickly sorted by clicking the column header (eg. project, job, or status), which will set the sort attribute and the sort order, clicking the header again will invert the sort order.

Clicking a row will select the associated job, and clicking it again will deselect it. The checkbox at the far left of the row shows whether the job is selected or not (and can also be clicked to toggle the selection on or off).

The table view (as with the other views) can be saved as a CSV by clicking the download button at the bottom right of the jobs footer. This will export the entire data selection with the same filter, sorting, and granularity options as set in the table view.

## Job Groups

Job groups are an organizational feature that allow you to collapse multiple jobs into a single card to maintain the legibility of a workspace. They are designed to compact branches of jobs that are not being used for active processing but are still required for reference or continued processing in the future.

Job groups are simple constructs that are meant to be created and deleted without much overhead. Options for a title and description allow you to clarify a group’s purpose, and can help to document the thought process behind branches of processing when coming back later.

### Creating Groups

To create a group you must select all of the jobs you would like to have included in it. From here you can select the “Group Jobs” option from the job sidebar actions menu, or right click on any of the selected jobs to open the quick actions menu and select the option from there. Alternatively you can use the keyboard shortcut `command` + `G` . All of these options will open the job group dialog with fields available for a title, description, and colour. All of these fields are optional, and if you would prefer to leave them blank a default title and colour will be applied. Clicking the “Create Group” button at the bottom of the dialog will create the group (This button is focused by default when the dialog opens, allowing the group to be created with all default fields by pressing the `enter` key).

<figure><img src="/files/2KPAYik5c7b5WCQdSfNN" alt=""><figcaption></figcaption></figure>

### Selecting Groups

Groups can be selected by clicking on the group card or by clicking the group menu button on the footer to open the group search menu and choose the group you would like to select. This will also navigate the view to the selected group.

Groups are meant to be lightweight, and because of this, they do not have any specific metadata associated with them. When selecting a group you are essentially just selecting the jobs that are in that group. This is reflected in the sidebar, which is identical to the multi-selection sidebar that is shown when selecting multiple jobs whether they are grouped or not.

<figure><img src="/files/J2SHDx4jGMNZt7hNfvgt" alt=""><figcaption></figcaption></figure>

### Expanding and Collapsing Groups

All groups, and jobs that belong to a group, will appear with a group widget on their card. This widget includes the group ID and a +/- button.

Clicking the +/- button will expand or collapse the group. Expanding the group will show all of the jobs contained in the group directly in the view, while collapsing the group will show only the group card. Keyboard shortcuts are also available for expanding and collapsing groups. When a group or group job is selected or hovered, you can press the `E` key to either expand or collapse it.

Clicking the group ID will select all of the jobs in the group, this allows you to select all grouped jobs whether they are collapsed into a group card or expanded as individual jobs.

The view footer also includes +/- toggle buttons beside the group count that allow you to expand or collapse all of the groups in the workspace.

<figure><img src="/files/4TBrnxYyf4TyNfsfdcc6" alt=""><figcaption></figcaption></figure>

### Adding Jobs to a Group

You can add a single job or multiple jobs to an existing group by selecting the jobs and then right clicking any of them to open the quick actions menu. From here you can navigate to the “Add to Group” or “Add {x} Jobs to Group” option and select the group you wish to add these jobs to from the submenu.

<figure><img src="/files/U8ljLfYj62bGXkjwIvvo" alt=""><figcaption></figcaption></figure>

### Removing Jobs from a Group

You can remove any number of jobs from an existing group by selecting the jobs and then opening the quick actions menu by right clicking on any of them. From here you can select the “Remove From Group” or “Remove Jobs from Group” option. Alternatively you can use the `command` + `shift` + `G` keyboard shortcut to quickly remove jobs from a group. If you remove all of the jobs from a group, the group itself will automatically be removed.

<figure><img src="/files/TKM7dX2ZJH157mOx9GQO" alt=""><figcaption></figcaption></figure>

### Removing a Group

A group can be removed by hovering the group card to show its action buttons on the bottom right hand corner of the card, and then clicking the ungroup button. **This will not delete the jobs within the group, only remove the group itself.**

<figure><img src="/files/dJkUPDzOClIFacdLM2HO" alt=""><figcaption></figcaption></figure>

### Navigating to a Group

Groups have all of the same navigation options as job cards. This means that a selected group or a selection of group jobs, can be automatically panned and zoomed to by clicking the target button on the filter bar (or by pressing the `T` key).

Groups can also be searched and found in the group menu to select and navigate to the group card or the expanded group job cards. This menu can be accessed by clicking the group menu button in the footer.

<figure><img src="/files/PXlEHYRfsRxbXtPSQDKT" alt=""><figcaption></figcaption></figure>

### Missing Connections

A Job Group is simply a set of jobs. However, not every set of jobs can be made into a Job Group. For a set of jobs to be a valid group, the following rule must be true: For any two jobs in the set, all the jobs on the chain connecting those two jobs must also be inside the set.

This rule is a necessary condition in order for the Job Group to be rendered as a single card in the tree view and is therefore a requirement.

As the simplest example of this rule, consider a job chain `A` → `B` → `C`.

The set `A,C` is not a valid group, because `B` is on the chain connecting `A` and `C`. However, `A,B`, `B,C`, or `A,B,C` are valid groups.

When creating a group, adding jobs to a group, or removing jobs from a group, if the rule above is violated, a `Missing Connections` error will occur. Sometimes, with a large group or a complex tree, it is not easy to tell which jobs are causing the error.

When you encounter this error, we advise that you carefully inspect the selection of jobs that you have made, and make sure that you have not missed any jobs that are on chains between jobs that are inside the group. It is important to make sure to check for, and link or move, any “dependent” jobs from other workspaces that are needed in order to make a valid group into the current workspace.

<figure><img src="/files/4XGY4ZIX6EpxaDy0R3MI" alt=""><figcaption><p>Missing connection error and resolution when creating a group</p></figcaption></figure>

<figure><img src="/files/J0RRYLB5wDfvui6c4hDu" alt=""><figcaption><p>Missing connection error and resolution when adding jobs to an existing group</p></figcaption></figure>

## Shortcuts

*Note: `command` key represents `command` on Mac and `ctrl` on Windows or Linux*

* `O` will toggle the outputs view on all cards.
* `E` will toggle the expanded/collapsed state of a group when selected or hovered (hover takes precedence).
* `T` will navigate the view to the currently selected job or set of jobs.
* `command` + `G` will group selected jobs.
* `command` + `shift` + `G` will ungroup selected jobs.


# Creating and Running Jobs

Working with the Job Builder, Job Cart and Job Quick Actions.

Jobs are the atomic processing unit within CryoSPARC. You progress through each stage of the cryo-EM data processing pipeline by building and connecting jobs together. Each job contains a set of inputs and parameters, and generates outputs to feed into other jobs.

Data processing workflows are often iterative and non-linear, and CryoSPARC jobs are designed to fit well within these workflows.

## Job Builder: New Job

The job builder is the primary method of creating new jobs in CryoSPARC. It is available through the ‘Builder’ tab on the sidebar.

{% hint style="info" %}
You can activate the job builder through the spotlight by typing `command` + `k` and typing “Job builder” or “New job”.
{% endhint %}

When the builder opens the search bar is automatically focused, making it easy to filter the list for a job you would like to build. While the search bar is in focus, you can use the up and down arrow keys to navigate through the list or skip to the top or bottom of the list with the shortcut `command` + `up/down`.

The new job list is organized into sections based on job category, and individual jobs are marked with helpful tags:

* **New**: this job has recently been introduced or upgraded
* **Beta**: this job is actively in development and may not be as stable as others
* **Import**: this job allows you to bring external data into CryoSPARC
* **Utility**: this job performs a useful action to augment or transform the inputs
* **Interactive**: this job spawns an process you can interact with to view and filter data
* **GPU**: this job is single GPU accelerated
* **Multi-GPU**: this job is GPU accelerated and can run on one or more GPUs
* **Live**: this job is used in CryoSPARC Live Sessions

You can further refine the list of jobs presented by expanding the filter options and selecting tags.

<figure><img src="/files/lEqP8ReQu3WJnX9ogiYT" alt=""><figcaption><p>The list of jobs reduced to only those which can run on multiple GPUs</p></figcaption></figure>

{% hint style="warning" %}
To view legacy jobs within the builder, expand the filter and toggle the ‘Show legacy jobs’ option. Keep in mind these jobs may not be stable or available in future versions of CryoSPARC.

![](/files/awXswdSaBlFK99ljynb5)\
Legacy jobs in CryoSPARC v4.0.0
{% endhint %}

## Job Cart

A new method of creating jobs in CryoSPARC v4 is through the job cart. Rather than choose from a list of jobs, the cart allows you to compile a set of job outputs that can be used as inputs for another job. The job cart automatically filters the list of new jobs based on their input requirements.

Job cart outputs can be populated from the job card outputs view, accessible in two ways. The first is from the job card header using the outputs action button. Clicking this will toggle the card view, hiding any images and information in the main area, and showing instead a list of the job’s outputs.

<figure><img src="/files/fa1fW49NeSiNqUZUj3Dh" alt=""><figcaption><p>A list of job cards with the outputs mode active. All output groups will be listed alongside their item count. Clicking any group will add it to the cart.</p></figcaption></figure>

The other way to activate this view is by holding down the `shift` key on your keyboard. This will toggle all job cards to the outputs view, and allow viewing and selecting these outputs as long as the key is held down.

Clicking on these output options will add them to the job cart. From here you have options to clear the selected outputs, modify the number of outputs of a type being used, and view jobs which have inputs matching the selected outputs. For each output that is added to the cart, a preview image and count will be displayed.

Similarly to the job list on the ‘Builder’ tab, you can use the arrow keys to navigate the list and press `enter` to create the job. Selecting a new job from the job cart will automatically create that job and connect all of its inputs with the available outputs.

<figure><img src="/files/TnkHC3HTlz83UTiA6et0" alt=""><figcaption><p>With an output group added to the cart, the list of applicable jobs is automatically filtered. You can further filter this list via the search input and hover over any job to view additional details. Once you click to select or press enter, the job will be created and all groups within the cart will automatically be connected to the new job.</p></figcaption></figure>

As projects are standalone containers of jobs, you can only add output groups from the same project into the cart. You must also be in a workspace of the relevant project to create a job from the cart.

## Job Quick Actions

In addition to the builder and job cart, you can quickly create jobs from a list of common actions available in the context menu of a job card. For example, a common action after running a 2D Classification job is to select certain classes via the Select 2D Classes interactive job. With job quick actions, you can simply right click on the job card and select ‘Queue Select 2D Classes’ to build and queue the interactive job without any manual intervention.

<figure><img src="/files/O6fUFtKTWwY1ZCZ9VyIe" alt=""><figcaption><p>Right-clicking on a 2D Classification job presents two quick actions. Clicking ‘Queue Select 2D Classes’ will automatically create and queue this interactive job.</p></figcaption></figure>

Quick actions are available for almost all job types in CryoSPARC. Most created jobs will enter building status and wait to be queued manually, while interactive jobs will automatically queue and run until they enter the waiting status (where user intervention is required).

Some jobs such as *Ab-initio Reconstruction* contain quick actions that allow for multiple jobs to be created, one for each class:

<figure><img src="/files/SEQAMzx3itldzuvQhOBA" alt=""><figcaption><p>Right-clicking a multi-class Ab-initio Reconstruction presents an option to build a refinement job for each class that was reconstructed.</p></figcaption></figure>

## Building a Job

When you create a new job, the “Builder” sidebar tab displays the building (or editing) view of a job. Here, you can connect inputs, configure the job’s parameters and queue it to be run.

Jobs active in the builder tab are sticky, meaning that you can continue to select and view other jobs while being able to add inputs and edit parameters of the building job. This allows you to view another job and drag and drop the output groups into the input groups of the building job.

To activate the job builder on any building job, press ‘b’ while the card is selected or click on the ‘Build’ button at the bottom of the card.

To stop building a particular job, press ‘b’ or click on the deselect button in the sidebar.

<figure><img src="/files/0PsH8WbZAdCs37HUPtsd" alt=""><figcaption><p>The job builder sidebar tab active with an Inspect Particle Picks job in building status</p></figcaption></figure>

### Inputs and Parameters

Enter any required parameters, and connect any required inputs. Inputs to a job are the data or files that will be processed by the job. For example, the input to a Patch CTF Estimation job are micrographs. Inputs for a job can come from another job inside the Workspace in which you are actively working, or any other Workspace within your current Project.

See [All Job Types in CryoSPARC](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc) for detailed information about the inputs, parameters and outputs for each job type.

### Drag and Drop to Connect

Connect the required input(s) by dragging and dropping the output(s) from another job.

<figure><img src="/files/2EslHCAp4mQTMESCasrN" alt=""><figcaption></figcaption></figure>

If you have successfully connected the output(s), you will see the job number appear below the input, as well as an option to Remove if you have accidentally connected the wrong output:

<figure><img src="/files/JhfYKMJAG7w8Sn0curW1" alt=""><figcaption></figcaption></figure>

For more details how how CryoSPARC inputs and outputs work, along with more advanced features of the job builder, refer to the following page: [Tutorial: Job Builder](/guides-for-v3/job-builder-tutorial)

### Default and Custom Parameters

Most job parameters in CryoSPARC have a default value. When you change a parameter to a value other than the default, it is marked as a custom parameter and will be highlighted in green. At any time, you can click the ‘Reset’ button to cycle between resetting the parameter to its default value or clearing the parameter value altogether.

At the top of the ‘Parameters’ section you can easily view how many parameters have been customized and toggle between displaying all, only defaults or only custom.

<figure><img src="/files/91N4ngGm0cAZhNPD94gI" alt=""><figcaption></figcaption></figure>

Each custom parameter is also listed directly on the job card, making it easier to distinguish many building jobs from one another.

<figure><img src="/files/XLaDMyX5oP9AMXJR6jjb" alt=""><figcaption></figcaption></figure>

### Viewing Advanced Parameters

By default, the job builder will display a set of parameters for a job that are more commonly edited. To view the full list of parameters available for a job, select the ‘Advanced’ toggle at the top of the ‘Parameters’ section:

### Copy and Paste Parameters

CryoSPARC v5 now includes the ability to copy a job's parameters to the clipboard and then paste them to another building job to populate its parameters.

<figure><img src="/files/Bn4PHx1rOIZhKCLeofeO" alt=""><figcaption></figcaption></figure>

Use `command`+`option`+`c` to copy parameters from a selected job and then select another job and use `command`+`option`+`v` to paste these parameters to it. Parameters can be copied from a job in any state, but pasting can only be used on jobs that are in building status; this will overwrite the job’s current parameters. Pasting can be undone by clicking on the "undo" button in the success notification that will pop up in the bottom left corner of the UI.

Parameters copied from a job of one type can also be pasted onto a job of a different type. This can be useful between job types with many of the same parameters, such as refinements. Any parameters not shared between the job types will simply not be applied.

Parameters that are copied using the keyboard shortcut will be saved to the clipboard as valid [blueprint](/application-guide/blueprints) JSON and can be pasted, modified, saved, and shared as valid [blueprints](/application-guide/blueprints). Parameters can also be copied from the job builder by using the `command`+`option`+`c` shortcut while focused on a specific blueprint, or by selecting the "Copy" option in the blueprint options menu.

There is also a button group present at the top of the parameters section in the job dialog and job sidebar parameters sections with two options to either copy as blueprint JSON (for pasting) or as simple key value text for reference. Switching the pre-existing "Custom/All" toggle will filter parameters by type, and enable either copying all parameters or only custom parameters.

<figure><img src="/files/n9EQkceewunSCQ5vpnVU" alt=""><figcaption></figcaption></figure>

## Queueing a Job

When you are satisfied with a job’s inputs and parameters, you can queue the job. Click 'Queue Job' on the Job Builder. This will bring up the Queue Slide-over.

1. Select the lane where the job should run. By clicking on the "Run on specific GPU" tab, you can specify a GPU or list of GPUs this job should use. Queueing a job in this way will launch the job immediately if possible, and bypass the scheduling system.
2. Optionally, you can set a title or description for the job and adjust the Job Priority. See: [Priority Queuing](/setup-configuration-and-management/software-system-guides/tutorial-priority-job-queuing).
3. Click 'Queue' to queue the job.

<figure><img src="/files/ZimfT1x67FgRzcoQ3X5i" alt="" width="369"><figcaption><p>The Queue Slide-over in CryoSPARC v4.0</p></figcaption></figure>

In version 5.0 and later, the Queue slide-over provides detailed visibility into the compute resources utilized on each node lane and target. It proactively displays a warning when attempting to queue jobs to a lane with resources that are already fully allocated, helping to prevent submission to over-utilized lanes. Jobs may also be queued directly to a specific target by selecting the radio button located to the left of the target card. It is important to note that the displayed utilization statistics reflect only the resources consumed by running CryoSPARC jobs and do not account for resources used by other programs or processes outside of CryoSPARC.

<figure><img src="/files/5gCq7ltIDWNw7yBAG5nB" alt="" width="375"><figcaption><p>The Queue Slide-over in CryoSPARC v5.0+</p></figcaption></figure>

<figure><img src="/files/rJ0fSUxSVYJiFavodypO" alt=""><figcaption><p>The Queue Job sub-menu in CryoSPARC v4.0</p></figcaption></figure>

Similar to the Queue Slide-over, the Queue Job sub-menu in v5.0+ surfaces resource utilization information to help you avoid queueing to over-utilized lanes:

<figure><img src="/files/SZukSycK237MOf6CTtw4" alt=""><figcaption><p>The Queue Job sub-menu in CryoSPARC v5.0+</p></figcaption></figure>

### Queuing Chains of Jobs that Run Automatically

You can queue a chain of jobs, which will each commence as soon as their respective required inputs (i.e., the outputs from jobs earlier in the series) become available.

For example, while an Import Movies job is running, open the job card, and drag and drop the `imported-movies` output into a new Patch Motion Correction job in the Job Builder. Queue the Patch Motion Correction job and it will appear in "Queued - waiting because inputs are not ready" state until the imported movies become available, at which point it will start running automatically.

See also, [Management and Monitoring](https://guide.cryosparc.com/setup-configuration-and-management/management-and-monitoring) if you wish to create and launch jobs through the command line interface.

## Running a Job

When a job is queued, it will be launched on a processing node or cluster. To view details of a job, select it and press the spacebar to inspect it.

{% hint style="info" %}
Queued jobs no longer automatically open like in CryoSPARC v3.
{% endhint %}

As a job runs, more information will display in the browse view. When viewing jobs in the cards view or tree view, a preview image of results and info tags will display on the relevant job card, providing a summary of results without having to inspect the job.

The next section covers more information on how to inspect job data and interpret results.


# Inspecting Job Data

## Selecting Jobs

The sidebar automatically displays information pertaining to the selected jobs. Click to select a job, use the arrow keys to navigate between jobs, and press `command` +`click` to add or remove jobs from your selection.

<figure><img src="/files/FTF4HSG8Dh6ExRyW6AEy" alt=""><figcaption><p>Selecting a single job will display its details in the sidebar ‘Details’ panel. You can press Spacebar, ‘View Job’ or click the job ID/type on the header of the card.</p></figcaption></figure>

## Job Info Tags

Almost all jobs include summary information presented on the cards, table, and tree view - these are called info tags. For example, the *Blob Picker* job card now displays key information in tags anchored around the preview images to display the minimum and maximum diameter configured, how many micrographs were processed, how many particles were picked, and the average number of particles picked per micrograph:

<figure><img src="/files/55pKMRjwU8X1GiqoIC3Y" alt=""><figcaption></figcaption></figure>

The combination of job preview images, info tags, and additional metadata displayed on the job card footer makes it easier to distinguish between multiple runs of the same job type with different parameters and results.

## Job Dialog

You can inspect a job to view much more information about it, including an interactive dashboard, real-time log of events, and interactive utilities. To inspect a job, select it and press spacebar or click on the ‘View Job’ button from the sidebar or click on the job ID and type on the header of the job card.

<figure><img src="/files/n11L2DPivX9aVoa55SDY" alt=""><figcaption><p>Inspecting a completed 2D Classification job presents the Event Log tab for quick inspection of the latest events.</p></figcaption></figure>

The job preview dialog opens above the current page and can be dismissed at any time by pressing the spacebar or escape key or clicking on the ‘x’ button.

<figure><img src="/files/hg1RSVFiJ8sGs88tMpM9" alt=""><figcaption><p>You can close the job inspection dialog by clicking on the ‘X’ button on the top-right of the dialog or press the <code>Spacebar</code> or <code>Escape</code> key</p></figcaption></figure>

At the top of the job inspection dialog is a breadcrumbs element depicting the hierarchy of a job. Clicking on a project or workspace will navigate to that page. Clicking on the job will navigate to the workspace page and scroll to that job.

Alongside the breadcrumbs, you can star the job and view the elapsed time.

<figure><img src="/files/8O8odqtnkkMzHi2xoSQ4" alt=""><figcaption></figcaption></figure>

The bottom of the dialog contains a footer with helpful information such as what processing node the job was queued on, how many GPUs were allocated and what custom parameters were set:

<figure><img src="/files/9a6hjBcSpamWrKWk6HIr" alt=""><figcaption><p>The job dialog footer contains metadata with additional details visible when hovered.</p></figcaption></figure>

The job inspection dialog is comprised of various tabs that provide more detail into the processing history and results of the job. To the right of the dialog’s main content area is a collapsible sidebar that lists all output groups the job produces. From this sidebar you can drag and drop output groups into the job builder or download results.

<figure><img src="/files/6EtP6xlJoRNBAs7VSOSH" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
Close the sidebar (top-right button or `command` + `/`) to expand the visible area of the job inspection dialog.
{% endhint %}

### Dashboard

The job dashboard is a new view introduced in CryoSPARC v5.0. It is the default tab when opening the job dialog and includes a concise overview of key job stats, errors and warnings, all chart images created by the job sorted by general relevance, and a collapsable panel which shows the most recent text events output by the job.

<figure><img src="/files/nAaW7Xzpmtprve6qJBer" alt=""><figcaption></figcaption></figure>

#### Info Cards

Info cards are located at the top of the panel. They include all of the data available in the [info tags](#job-info-tags) that appear on the job card as well as as some extended information. Values in the info cards can be clicked to copy to the clipboard.

<figure><img src="/files/MQ4aeKxAAPRG6DPOn16K" alt=""><figcaption></figcaption></figure>

#### Errors and Warnings

Errors and warnings are embedded directly into the dashboard and organized together to quickly surface this information.

<figure><img src="/files/HJP76PB2KvDUWBkIuI3H" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/SFvvWw5C4ruW2d9ovQdS" alt=""><figcaption></figcaption></figure>

Errors/warnings populate the module in descending order and can be copied to the clipboard by clicking the “copy” button in the top right corner of the message (becomes visible when hovering the message).

#### Charts

The charts section contains all of the chart images that have been output by the job. These are the same images available in the job’s event log, ordered by relevance, and set to the most recent iteration by default.

<figure><img src="/files/TvYJARPEXhCDbgPdgvOM" alt=""><figcaption></figcaption></figure>

#### Expanding and Collapsing Charts

Charts exist in collapsable sections that correspond to their “chart type” (eg. Real Space Slices, 2D Classes, etc). Charts are sorted by general relevance with the least relevant charts sorted to the bottom and collapsed by default. The expanded/collapsed state for charts can be controlled in a few ways:

* A chart can be expanded or collapsed by clicking on its individual section header.
* All charts can be expanded at once by clicking the “Expand All Charts” toggle button (arrows pointing away from a dotted line) on the Charts section header.
* All charts can be collapsed at once by using the “Collapse All Charts” toggle button (arrows pointing towards a dotted line) on the Charts section header. This can be helpful when trying to find a specific chart type in a job that outputs many charts.
* Charts can be reset to their default expanded/collapsed state by using the “Reset” button on the Charts section header.
* All charts can also be expanded or collapsed by holding down the `command` key and clicking the header of an individual chart type.

#### **Iterating Charts**

Charts are set to the most recent iteration available by default (if applicable, for jobs that output multiple iterations of the same chart type as they progress). By moving the slider, iterations can be scanned to see how the job progressed or is currently progressing. There are two iteration sliders available in the view:

* The **global iteration slider** is available on the Charts section header. This slider will progress through all available iterations across all charts.
  * Many charts do not output new images on each iteration as the changes are not significant enough to warrant it. If a chart does not have an image for the iteration specified in the global slider, it will continue to display the chart image for the last available iteration and will show a small orange “unlinked” icon in its header to indicate that the chart image iteration does not match the slider’s current iteration.
* The **local iteration slider** is available in the header of a specific chart. This slider controls the iteration for that chart only.

#### Chart Lightbox

The chart lightbox allows a chart image to be expanded to its full size and viewed in isolation.

Open the lightbox by simply clicking on any chart image on the dashboard. The lightbox can be closed by clicking outside of the main area on the darkened backdrop, clicking the red “X” button in the top right corner of the content area, or pressing the `escape` key.

<figure><img src="/files/kUosobNeMT8NVgZcAzht" alt=""><figcaption></figcaption></figure>

The lightbox includes a number of features to help you navigate between chart images quickly and efficiently:

* The **Chart Type Switcher** is the second button group on the left side of the header. It has a central button which displays the chart type (eg. Real Space Slices, FSC, etc). This button opens a menu containing all chart types the job generates, allowing you to switch the lightbox to any chart type you wish to view. The arrow keys on either side will iterate incrementally through all available chart types. Both of these actions will reset the iteration and index between chart types.
  * Chart types can also be switched between by holding down the `command` key and pressing either the `leftarrow` or `rightarrow` keys to navigate.
* The **Chart Switcher** is located in the bottom left corner of the lightbox. It includes navigation arrows to iterate between charts, an input to manually set the chart index, and a button that indicates the number of total charts available for the chart type (clicking this button selects and switches to the final chart of that type).
  * Charts can be also be switched between by pressing either the `leftarrow` or `rightarrow` keys to navigate.

The chart’s iteration can be controlled using the slider located on the right side of the lightbox header (if applicable). This slider operates the same way as the **local iteration slider**, and will update the current iteration for the specific chart being viewed.

The lightbox also includes a click to copy title for the current job image directly below the header, and a click to copy timestamp representing when the chart was created by the job in the bottom left corner of the footer.

The download module in the centre of the footer includes buttons that can be clicked to download any of the available file types for the current chart. In the case where a chart type contains multiple iterations, a GIF download option is presented.

#### Interactive Charts

Interactive event charts are embedded directly into the job dashboard if they are available. These charts allow for deeper exploration of generated output data through rich interactive features.

Interactive charts are contained in a collapsable "Interactive" section module. Each chart is contained in its own named section which can also be collapsed.

Above the chart is a control bar where all relevant filters and chart settings are shown. Below the chart is a footer which includes timing information and a module to download event data files (eg. a png of the event image).

{% hint style="info" %}
Interactive charts are currently only available for 3D Variability Display jobs
{% endhint %}

<figure><img src="/files/JqtbrjhuLF33xaxmomDh" alt=""><figcaption></figcaption></figure>

#### Embedded Event Log

The job dashboard includes an expandable text only event log attached to the bottom of the panel. This event log includes the 120 most recent text events (including errors and warnings) and tails the log, always showing the most recent events as they appear from the bottom.

<figure><img src="/files/ovIWPItPHKGQTVf7SbWz" alt=""><figcaption></figcaption></figure>

The event log is pinned to the dashboard footer and will always be visible when scrolling. The event log panel can be expanded or collapsed by clicking the +/- toggle button. Its height can also be adjusted by grabbing the top edge of the header (which will become highlighted blue) and then dragging it to the desired height.

The event log also includes a button group in its header to download relevant logs. The lefthand button downloads a pdf of the entire event log, while the righthand button opens a menu with the option to download the full job report.

### Event Log

As a job runs, events are published in real-time as a job processes. Events are viewable when a job completes. Above the log, a list of controls are available:

<figure><img src="/files/NKiOXfWuMN21k1loqRlQ" alt=""><figcaption></figcaption></figure>

* **Show from top**: Jump to the top of the event log, scroll to load events in chronological order.
* **Follow latest**: Skip to the last available checkpoint and listen for new events.
* **Select a checkpoint**: View a list of checkpoints and timestamps, click to jump to the start of the checkpoint.
* **Filter types**: Display events of a certain type (such as text or image).
* **Filter flags**: Display events of a certain job-dependent flag (such as central slices or FSC curves).
* **Show CPU usage**: Display CPU memory usage of the processing node alongside each log.
* **Show timestamps**: Display a timestamp alongside each log (you can also hover over the icon on the left side of each log to view the timestamp).

### Interactive

The interactive tab is only visible when running an interactive job that is in ‘waiting’ status. Please refer to the [Interactive Jobs guide](/application-guide/interactive-jobs) to learn more.

### Inputs and Parameters

In addition to the details sidebar, this tab displays all inputs and parameters that were configured. Select 'Show slots' to view low-level slots contained within each input group. Parameters will be marked as either default (grey/blue) or custom (green). Advanced parameters are denoted with an ‘A’. You can toggle between listing only custom parameters or all parameters.

<figure><img src="/files/70Iq3wTYasEsObezBPWh" alt=""><figcaption></figcaption></figure>

### Outputs

In addition to the output groups sidebar panel, the outputs tab displays a comprehensive overview of all the output groups and individual outputs a job creates. Please refer to the [Job Builder Tutorial](https://guide.cryosparc.com/processing-data/tutorials-and-case-studies/job-builder-tutorial) to learn more about how you can take advantage of the low-level results CryoSPARC generates.

<figure><img src="/files/cgkUGaGoBx0qkq9URhBW" alt=""><figcaption></figcaption></figure>

### Volumes

This tab is only visible when a job outputs one or more volumes. By selecting a volume from the left sidebar, it will load in the integrated volume viewer. You can set a threshold, zoom and pan, start or pause animation, and download the volume for inspection in an external software.

<figure><img src="/files/iSFDqdXjJoM4ruFRKZQV" alt=""><figcaption></figcaption></figure>

For volumes that also have an associated colour map (such as running a refined volume through [Local Resolution Estimation](/processing-data/all-job-types-in-cryosparc/post-processing/job-local-resolution-estimation)), an additional option will become available in the sidebar to view the coloured volume.

<figure><img src="/files/hUOPeURZJ6N9BxgjmXKI" alt=""><figcaption></figcaption></figure>

### Metadata

Useful for when archiving or debugging, the metadata tab contains all job data stored in the database. You can search for specific fields using the search bar. The path to the job logs location on disk is listed as well and can be clicked to copy.

<figure><img src="/files/bMvfpekYyUoW5OV5PNWa" alt=""><figcaption></figcaption></figure>

In addition to the job data, you can view the text log of the job by clicking on the ‘Log’ sub-tab:

<figure><img src="/files/NbegCEuqBwhKCZ1LJ6qT" alt=""><figcaption></figcaption></figure>

## Related Jobs

When a job is selected or inspected, the sidebar details panel will list an overview of what other jobs are connected to it via parent or children connections, or it being cloned from another job of the same type. Hovering over each job will display a tooltip containing an outline of that job, including what parameters it has been configured with. Clicking on a related job will inspect it. To open the related job in a new tab, click it while pressing the command (macOS) or control (Windows, Linux) key.

<figure><img src="/files/Nvj8Xi1mQrMGN2DAZqvs" alt=""><figcaption><p>Viewing a selection of related jobs connected to a Homogeneous Refinement</p></figcaption></figure>

<figure><img src="/files/0F5hz3RHwHKVmnn6P3Ee" alt=""><figcaption><p>Each related job will display key information about its configuration and outputs</p></figcaption></figure>

## Comparing Jobs

The comparison view is designed to enable analysis of multiple jobs side by side to compare and contrast their settings and results. It includes the ability to view differing parameters aligned between jobs, inspect job chart output images to observe how job progression has diverged during processing, and the ability to download outputs from multiple jobs in a single click.

<figure><img src="/files/ObH2D0mCrLg37jXbb7Zl" alt=""><figcaption></figcaption></figure>

#### Opening the Comparison View

The comparison view can be opened by selecting multiple jobs (`command` + `click` ) and then either clicking the “Compare” button at the bottom of the multi-select sidebar, or simply pressing the `spacebar`.

#### Header

The header includes all essential identifying information for each job, and is pinned to the top of the dialog to keep it visible while scrolling down to different sections.

<figure><img src="/files/QqNaVtdkHdGuzkIivngU" alt=""><figcaption></figcaption></figure>

This section contains the job’s unique ID, its job type, a button to star the job, and the job’s title (if applicable):

* The **job ID** has a dual function as a button that can be clicked to navigate to that specific job’s dialog. Using the web browser’s “back” button will navigate back to the comparison view.
* The **star button** allows the job to be starred directly from the comparison view, it will turn yellow when the job is starred.
* The **job title** field will be automatically shown if any job in the comparison has a title, and hidden if no jobs in the comparison have a title. The field can be manually shown or hidden by clicking the toggle button on the far right side of the Summary section header below the title. Clicking the title will convert it into an editable field, clicking away from it will save the title and switch out of the editing mode.

#### Summary

The summary section includes the job’s cover image, info tags, and description.

<figure><img src="/files/jtZzYzdVhhSlu5Af1Bqc" alt=""><figcaption></figcaption></figure>

Located on the far right of the summary section header are toggle buttons to show/hide the title and/or description as well as a button to download all relevant outputs from each job in the view.

* The **Download Job Outputs** button opens a sub-menu where the type of outputs that you would like to download can be checked or unchecked to add them to the download. The download can be made as individual files which will be downloaded sequentially, or a single zip. We recommend using the individual option generally, as the zip option can consume a significant amount of memory for large files, which can cause the browser to lock up or crash in some cases.
* The **job description** field will be automatically shown if any job in the comparison has a description, and hidden if no jobs in the comparison have a description. The field can be manually shown or hidden by clicking the toggle button on the far right side of the Summary section header. Clicking the description will convert it into an editor, and clicking the red “X” button in the top right corner of the editor will switch out of editing mode.

#### Parameters

<figure><img src="/files/gzbAwuDfti6hyfSBAWYF" alt=""><figcaption></figcaption></figure>

The parameters section allows for comparing and contrasting the differing parameters between jobs in the comparison view. The default setting is to show all parameters that are different between jobs (where at least one job has a custom value for that parameter) and to show all advanced parameters.

Each parameter is sorted and aligned in a grid with rows that stretch across all job columns. Custom parameters are shown in green with a corresponding “C” icon. An “A” icon is displayed if the parameter is advanced, and a “D” icon will be shown if it has a default value.

The Parameters header includes a button group where each button is a toggle used to hide or show sections for different, custom, and/or default parameters. A separate toggle button for “advanced” controls whether advanced parameters are shown or hidden across all parameter sections.

#### Charts

The charts section contains all of the chart images that have been output by each job. These are the same images available in the job’s event log, ordered by relevance, and set to the most recent iteration by default.

<figure><img src="/files/bzfMNazNoAc7lkSKrTEA" alt=""><figcaption></figcaption></figure>

#### **Expanding and Collapsing Charts**

Charts exist in collapsable sections that correspond to their “chart type” (eg. Real Space Slices, 2D Classes, etc). Charts are sorted by general relevance with the least relevant charts sorted to the bottom and collapsed by default. The expanded/collapsed state for charts can be controlled in a few ways.

* A chart can be expanded or collapsed by clicking on its individual section header, this will expand or collapse the entire row of charts for each jobs.
* All charts can be expanded at once by clicking the “Expand All Charts” toggle button (arrows pointing away from a dotted line) on the Charts section header.
* All charts can be collapsed at once by using the “Collapse All Charts” toggle button (arrows pointing towards a dotted line) on the Charts section header. This can be helpful when trying to find a specific chart type in a job that outputs many charts.
* Charts can be reset to their default expanded/collapsed state by using the “Reset” button on the Charts section header.
* All charts can also be expanded or collapsed by holding down the `command` key and clicking the header of an individual chart type.

#### **Iterating Charts**

Charts are set to the most recent iteration available by default (if applicable, for jobs that output multiple iterations of the same chart type as they progress). This can be controlled using the provided iteration sliders to dynamically switch between iterations. By moving the slider, iterations can be scanned to see how the job progressed or is currently progressing. There are two iteration sliders available in the view:

* The **global iteration slider** is available on the Charts section header. This slider will progress through all available iterations across all charts for all jobs in the comparison view.
  * Many charts do not output new images on each iteration as the changes are not significant enough to warrant it. If a chart does not have an image for the iteration specified in the global slider, it will continue to display the chart image for the last available iteration and will show a small indigo “unlinked” icon in its header to indicate that the chart image iteration does not match the slider’s current iteration.
* The **local iteration slider** is available in the header of each individual job column and controls the iteration of all of the charts for that job.
  * If any of the charts do not have a chart image for the specified iteration, a violet “unlinked” icon will appear in its header to indicate that the current chart iteration does not match the local iteration slider’s current iteration.

#### Chart Lightbox

The chart lightbox allows a chart image to be expanded to its full size and viewed in isolation.

Open the lightbox by simply clicking on any chart image in the comparison view. The lightbox can be closed by clicking outside of the main area on the darkened backdrop, clicking the red “X” button in the top right corner of the content area, or pressing the `escape` key.

<figure><img src="/files/n2z7XVCNurta1k1mjGUn" alt=""><figcaption></figcaption></figure>

The lightbox includes a number of features to help you navigate between chart images quickly and efficiently:

* The **Job Switcher** is a button group located in the top leftmost position of the lightbox header. It has a central button that displays the job’s project ID and job ID. When clicked the button opens a menu that allows you to select any job in the comparison view and jump to it. The arrow keys on either side will iterate incrementally through each job in the comparison view. Both of these actions will retain the chart type and index that is currently selected (if applicable) between jobs.
  * Jobs can also be switched between by holding down the `shift` key and pressing either the `leftarrow` or `rightarrow` keys to navigate.
* The **Chart Type Switcher** is the second button group on the left side of the header. It has a central button which displays the chart type (eg. Real Space Slices, FSC, etc). This button opens a menu containing all chart types the job generates, allowing you to switch the lightbox to any chart type you wish to view. The arrow keys on either side will iterate incrementally through all available chart types. Both of these actions will reset the iteration and index between chart types.
  * Chart types can also be switched between by holding down the `command` key and pressing either the `leftarrow` or `rightarrow` keys to navigate.
* The **Chart Switcher** is located in the bottom left corner of the lightbox. It includes navigation arrows to iterate between charts, an input to manually set the chart index, and a button that indicates the number of total charts available for the chart type (clicking this button selects and switches to the final chart of that type).
  * Charts can be also be switched between by pressing either the `leftarrow` or `rightarrow` keys to navigate.

The chart’s iteration can be controlled using the slider located on the right side of the lightbox header (if applicable). This slider operates the same way as the **local iteration slider**, and will update the current iteration for all charts in the job.

The lightbox also includes a click to copy title for the current job image directly below the header, and a click to copy timestamp representing when the chart was created by the job in the bottom left corner in the footer.

The download module in the centre of the footer includes buttons that can be clicked to download any of the available file types for the current chart.


# Low Level Results Interface

{% hint style="info" %}
The low level results interface is an advanced UI feature. It is only necessary when a workflow requires combining results from several different jobs.
{% endhint %}

## Outputs and Inputs in CryoSPARC

When a CryoSPARC job completes, it produces a number of *output groups*. These are familiar to all CryoSPARC users: output groups have names like Particles, Volumes, Exposures, etc. These outputs contain all of the information produced by the job for each kind of output item. Within output groups, there is an additional organizational level, called *output results*. Each output result has a name (e.g., `blob`, `ctf`, `locations`, `mscope_params`, etc.) and contains information of only a specific type. Together, a set of output results comprise an output group.

It is easier to understand groups and results with a concrete example. Consider a Particles output group. At the bare minimum, a Particles output must have all of the information about the particle images themselves: where to find them on the disk, their pixel size, etc. However, there is often more information than just the images. Perhaps the particles have CTF estimates or 2D or 3D pose estimates. Thus, the Particles output group contains each of these types of information as individual output results.

![](/files/2j5FZNr8Gz8n89HuNXPC)

When creating a new job, you connect the parent job’s output group to the descendent job (from the point of view of the descendent job, it’s an input group). Taking a concrete example, you might connect the Particles and Volume outputs from a Homogeneous Refinement to the relevant Local Refinement inputs. The Local Refinement now has access to all of the output results produced by the Homogeneous Refinement. As the Local Refinement proceeds, it will update the particles’ poses (in the alignments3D output result) and then produce its own output group with the updated alignments3D result.

![](/files/ki0eWT5f600BNMTAMEe0)

## When is the Low Level Results Interface useful?

In some cases, you may want to replace *only a single result in an output group* with the result from another job. For example, perhaps the Homogeneous Refinement was performed with a curated set of downsampled particles, but you want to perform a Local Refinement with the full size particle images. You could do this using standard CryoSPARC jobs:

1. Use Particle Sets Tools to find the overlap between the curated, downsampled particles and the uncurated, full-size particles.
2. Perform a Homogeneous Refinement of the full-size particles against the downsampled map to assign poses to the full-size particles.
3. Perform the Local Refinement.

This process would work, but consider step 2. We already have pose estimates for these particles, but they were found using the downsampled particles, so we cannot use them with the Extract from Micrographs output group using the normal job building interface. We are therefore forced to perform an alignment which is otherwise unnecessary and potentially slow.

If we could use all of the information from the downsampled particles *except for the blob output result* (which contains the particle image data at a specific box size), we could use the existing poses (the alignments3D result from the downsampled Homogeneous Refinement’s particles output group) with the full-size images (the blob result from the Extract from Micrographs job’s Particles output group). This output result swapping is performed with the low level results interface.

![](/files/Xt7S2OUBwUn22hzpaGHY)

## Using the Low Level Results Interface

{% hint style="info" %}
Data in this section comes from [EMPIAR-10833](https://www.ebi.ac.uk/empiar/EMPIAR-10833/), originally collected by Kim and colleagues.
{% endhint %}

In this section, we demonstrate the use of the low level results interface by working through the example above. Keep in mind that any result can be replaced using the low level interface — there is nothing unique about the blob output result which we replace here.

We begin by setting up an Extract From Micrographs job to extract particles picked with Blob Picker. In addition to extracting a full-size image, we use the `Second (small) F-crop box size (pix)` parameter to extract a much smaller image of each particle. These images will be used during particle curation, for which resolution is much less important than speed.

| Parameter                             | Value |
| ------------------------------------- | ----- |
| Extraction box size (pix)             | 452   |
| Second (small) F-crop box size (pix)  | 96    |
| Save results in 16-bit floating point | On    |

We proceed through a standard particle curation pipeline using the 96 pixel images until we have a clean particle stack. Because these particle images were downsampled to a pixel size of approximately 5.25 Å, a Homogeneous Refinement easily reaches the Nyquist resolution of 10.5 Å.

![](/files/hdrv97u18YJ5nFsxbabB)

We’d like to perform a Local Refinement of the full-size particle images using the poses found with this Homogeneous Refinement using the low level interface. To do so, we will first create a Local Refinement job and connect the Particles and Volumes output groups to the Local Refinement as usual.

![](/files/ZAGQKfNZNsXAg6a7w7Qv)

We can reveal an input or output’s results by clicking the small arrow on the left-hand side of the group’s box:

![](/files/JZtNJhJhhDx0phhjjysO)

<details>

<summary>What is a passthrough?</summary>

In the figure above, you may note that the particles’ location and pick stats are *passthroughs*. When jobs do not modify the information in an output group, they are treated as passthroughs. Under the hood, this means that CryoSPARC does not save a new copy of the passed-through information for the new job. Instead, the job’s output contains the location of the original information. More information about passthroughs is available [here](https://guide.cryosparc.com/guides-for-v3/job-builder-tutorial#passthrough-results).

CryoSPARC and CryoSPARC Tools both automatically fetch passthrough information, so in most cases you can treat passthroughs as if they were a typical output group slot. External programs which directly read `.cs` files, however, will not have access to passthroughs without special attention. To create a `.cs` file which contains all of the metadata (including passthroughs), [export](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/tutorial-data-management-in-cryosparc#use-case-share-a-particular-job-with-another-user) the job first.

</details>

These boxes display which specific output results are connected to the input group of the currently building job. As mentioned above, they are automatically connected with the appropriate output results from the output group. However, we want to replace the blob result from the Homogeneous Refinement with the blob result from the full size extracted images. To do so, we open the Extract From Micrographs job and navigate to the Outputs tab. This tab contains a larger version of the output groups seen in the sidebar. Within these larger output groups, we see small boxes corresponding to the individual output results.

![](/files/5bHRisEawAJafDOeTZCv)

To replace the Homogeneous Refinement’s blob result with that of Extract from Micrographs, we drag *the blob output result* from the Particles Extracted output group to the blob result of the Particle stacks input group.

![](/files/0t4A5T8zkuRHqlghJZBd)

Note that while dragging an output result, any result slot which can accept that result turns green. After dropping the result into the slot, the job number displayed in the input result updates to reflect the change. In this case, the result initially reads `J688.particles.blob.F`. This indicates that the result uses the `blob` result of the `particles` output group from the final iteration (this is the meaning of the `F`) of job `688`.

Once the result is replaced using the low level results interface, it reads `J683.particles.blob.F`, meaning it will use the `blob` result from `J683`'s particles group. The other slots are unchanged, so the `ctf` and `poses` from `J688` will still be used.

If we run this Local Refinement, we see that the initial poses from the Homogeneous Refinement are used, but the full-size images produce a much higher resolution map.

![](/files/cNjX9d9ppxnEqvG17Bq7)

## Using Output Results from Earlier Iterations

Some jobs save several versions of output groups. By default, the final version produced by the job is used. However, in some cases it may be desirable to use the output group from an earlier iteration. For example, early iterations of Ab-Initio Reconstruction are often good sources of junk volumes for particle curation. The Low Level Results Interface can also be used to select a specific version of an output result.

You can access earlier versions by clicking the `Versions` dropdown at the bottom of the output group box. This opens a panel listing each of the iterations. Click an iteration to select it. The name of the output result will update to reflect the change — for example, going from `J683.particles.alignments3D` to `J683.particles.alignments3D.3` if Iteration 3 is selected. Connecting the output result to an input’s result slot will now use the selected version.

<figure><img src="/files/gFmL8FjpQmwaRRKXStKW" alt=""><figcaption></figcaption></figure>


# Filters and Sorting

Cryo-EM data processing can become complex quickly. A large number of jobs, workspaces, and projects will be created in typical use. It therefore becomes very important to be able to filter and sort through items in CryoSPARC.

## Filter System

When browsing CryoSPARC, the filter and sorting options are displayed in the control bar at the top of the main content area:

<figure><img src="/files/rIVIezfbpDGma35OFeZw" alt=""><figcaption></figcaption></figure>

### Quick Filter Buttons

<figure><img src="/files/yqpKwlPIgI3HFXg4vRbn" alt=""><figcaption></figcaption></figure>

Quick filter buttons help quickly filter the displayed items.

#### **General Quick Filter Buttons**

* **Starred:** This button will toggle on the starred filter causing only starred items to be shown. This allows you to quickly locate the your important projects, workspaces, sessions, or jobs.
* **Mine:** This filter will show only items created by you when toggled on. This can be useful if you are an admin user, or simply to filter out items shared with you to find your own data quickly and easily.
* **Tags:** This filter will show all applicable tags with the granularity tags (eg. project or workspace specific tags) shown first, and general tags shown before. The numbers on the right side of each tag entry represent the number of items in the instance that have that tag applied.

#### **Specific Quick Filter Buttons**

* **Sessions**
  * **Status:** The session status filter allows you to filter all sessions by running, completed, or paused. These filters are additive and adding more will show jobs of all selected status types.
* **Jobs**
  * **Status**: The job status filter allows you to filter jobs by any available status, or combination of statuses. These include: Building, Queued, Launched, Started, Waiting, Running, Killed, Completed, and/or Failed.

### Filter Bar

The filter bar is where more advanced filter compositions can be created and managed. It contains all possible filter options. In order to access the filter bar, you must click the “Filters” button on the control bar to open the underslung module.

<figure><img src="/files/THlJINyO6pum7z6a594g" alt=""><figcaption></figcaption></figure>

#### Overview

The filter bar is a powerful system for adding, composing, and managing filters. The filter bar has separate filter options relevant to the type of item being filtered.

{% hint style="info" %}
Filters that you apply are stored in the browser URL bar. This means that you can bookmark or copy the URL from the browser and return to any filtered view that you have created. You can also share the URL with other users who have access to the same projects, and they will see the same filtered view that you have created.
{% endhint %}

#### Walkthrough

The filter bar is composed of three parts: the filter button, input, and clear button. Clicking the filter button or in input area will open the filter options menu. In the input area, you can also type a word or phrase in order to quickly chose an attribute by which to filter. The clear button allows you to clear all of the filters in the bar in one action, reseting the filter state and showing all available items in the view below.

<figure><img src="/files/XFLt8eL2NNkUS867BOQ5" alt=""><figcaption></figcaption></figure>

After opening the filter options menu, you can enter the menu by moving the cursor over it or pressing the down arrow on your keyboard.

There are three classes of filters: singular, compound, and toggle filters. Singular filters such as user filters have multiple options, but can only have one option active at a time. Compound filters such as tags, job type, or status, allow you to select multiple options and will additively filter results (eg. adding multiple status filter options, such as building, queued, and completed, will show all jobs that have any of those statuses).

Let’s begin by adding a building job status filter to see the resulting filter bar composition. After selecting the building status, the menu will remain open in order to allow you to add other status filters more easily. Click outside of the menu or press the escape key to close it.

<figure><img src="/files/6xmYDyj93S868Ovwi4DB" alt=""><figcaption></figcaption></figure>

Now that the filter has been added, it will appear in the input section of the filter bar. The light grey grouping indicator on the lefthand side of the filter tag shows the type of the filter added, in this case “Status”. The coloured filter tag on the righthand side shows the filter option that has been added, in this case “building”. By hovering over the filter tag a small red “X” indicator will show up on the top left corner of the tag, when clicked this will clear that particular filter option. Alternatively, hovering the grey filter grouping indicator on the lefthand side will also show a red “X” indicator on its top left corner, clicking this will clear all filters within that group.

Reopen the filter menu by clicking on the filter bar input section again (anywhere to the right of the currently existing tags). Add the status “killed” from the statuses submenu. This will add a second filter option to the statuses grouping.

<figure><img src="/files/RkAj6Cg0G7hbPCw2o8a5" alt=""><figcaption></figcaption></figure>

Now all jobs shown will either be in building status or killed status. Additional filters from different groupings can be added, such as a date filter, or job type filter, to further isolate results.

Filters can be added and removed to compose a variety of specific views allowing you to find jobs quickly and easily, or compose sets of results outside of the constraints of project or workspace groupings. This can be especially powerful when adding tags or starred filters, which operate as higher level selection groupings similar to upper level granularity groupings but with more flexibility.

### Sorting

Sorting options are located in the sort toggle button on the control bar beside the filter button. The sort button is composed of two parts, the sort order toggle on the lefthand side allows you to toggle ascending or descending order, and the sort option menu on the righthand side allows you to select what option you would like to sort by. By default, items are sorted by the “Date created” option. Projects and workspaces are sorted by default in descending order (showing the newest items at the top), while jobs are sorted in ascending order (showing the newest jobs at the bottom).

{% hint style="info" %}
Like filters, sort options are stored in the browser URL, and therefore are retained when bookmarking or copying/sharing the URL.
{% endhint %}


# View Options

Configurable view options allow the specific information displayed for items at each granularity and view type to be changed to suit your specific preferences. View options are available for nearly all information displayed apart from distinguishing content such as IDs and titles.

The view options menu can be accessed by clicking the “eye” icon button on the far lefthand side of the control bar at any granularity level. This button will open the view options menu with the current view mode section open and others collapsed (other view mode options, ie. card or table, can be accessed but are closed by default for clarity).

<figure><img src="/files/dh5f7mqVOJcfYp4lwQmG" alt=""><figcaption><p>View options panel for project cards</p></figcaption></figure>

Each view option in the menu is mapped to a corresponding item in the interface which can be displayed or hidden by checking or unchecking the option’s checkbox. For example, in the project cards view unchecking the “Tags” option will hide the tags on all project cards.

View options are scoped to each granularity and each view (eg. project cards and project table views are controlled by different view options so hiding the created attribute for cards will not hide it for the table), this allows very granular control of information to make each view customizable independently.

The table view includes a further option for customizability with drag and drop re-ordering of the table columns.

<figure><img src="/files/OfBmBeNocRUPZC2vYixe" alt=""><figcaption><p>View options panel for project table</p></figcaption></figure>

Each view option in the menu under the table view section can be dragged and dropped within the list to reorder the options. The order of the menu, top to bottom, corresponds to the order of the table columns, left to right.

All changes to the view options are saved to the database on a per user basis. This means that they will persist between sessions until changed by the user.


# Tags

Tags enable a useful level of user-customizable organization in your CryoSPARC instance. Tags are shared among all users within the instance and can be customized with a title, description, and colour. There are two categories of tags, a 'general' type tag that can be applied to items of any type, and item-specific tags that only apply to either project, workspaces, sessions, or jobs.

The main purpose of tags are to allow you to quickly find and view specific subsets of data either inside of a container (projects or workspaces) or across the entire instance. This gives a lot of flexibility for creating arbitrary groupings of items that you want to be able to find again quickly or compare and reference.

## **Accessing Tags**

The main access point for creating and navigating tags is the Quick Access Menu that can be expanded from the navigation bar on the left side of the app window. By clicking the “Tags” tab at the top of the menu you will be able to see all of the tags currently created in your instance. Each tag grouping is represented here in a collapsable drawer with the type of tag and total count of tags with that type shown on the drawer header. You can use the search bar at the top to filter the available tags and quickly find one that you are looking for. Beside the search bar there is a “+” button that will open the tag creation slide-over and allow you to create a new tag. Once created that new tag will be available to view and use. Each tag row in the menu shows the unique tag ID (e.g., T5, T20, etc.), the tag title, and a count of how many items have been given that tag (eg. a project tag of EMPIAR with a count of 10 has been assigned to 10 projects in the instance). Clicking on one of these rows will navigate you to the relevant view with a filter for that tag applied (in the case of general tags, a context menu will open when the row is clicked and allow you to select the type of item you want to view).

<figure><img src="/files/oWPHb1rKAcmRQFEjT3iK" alt=""><figcaption></figcaption></figure>

## **Applying Tags**

Once tag(s) exist in the instance, they can be applied to any relevant item. There are multiple ways to apply a tag depending on how you are interacting with the item. Let’s look at applying a tag to a project as an example.

By navigating to the browse section we will see all of our projects in the cards view. We can choose a project to apply our tag to and use either the quick actions menu, or the sidebar to add it. Right clicking the card or clicking the triple dot menu in the header will open the quick actions menu, we can then navigate down to the “Edit tags” menu item and into the sub menu with a list of relevant tags. Project tags will appear first and general tags below. Each tag in the menu has the same information as the rows in the quick access menu, the ID, title, and a count of how many items that tag has been applied to. By clicking any tag in the menu (lets take our EMPIAR tag for example) that tag will be applied to the project. Opening the menu again and viewing the “Edit tags” submenu will show a checkmark on any tags that have been applied to the card (in this case EMPIAR). Clicking the tag row again will remove the tag from the project.

<figure><img src="/files/0UVeooJckpHIlI5mHIPJ" alt=""><figcaption></figcaption></figure>

Tags can also be seen in the item sidebar, inside the “Details” panel (which is the first panel from the top). The tags row is just below the title row and displays all tags applied to the item. When this row is hovered, an edit button will appear, clicking this button will open the same “Edit tags” menu with the exact same functionality as the one available from the quick actions menu, discussed above.

Tags can be applied to projects, workspaces, sessions, and jobs in exactly the same way. The only difference is the granularity specific tags available.

## **Using Tags**

As mentioned above, the main use for tags is organizing data into subsets that either compliment the existing project and workspace demarcations, or acts as a wider aggregator around them. Tags are fundamentally custom filters, and as such operate functionally within the filter system. The control bar has a “Tags” quick filter button above the main content area, where tag filters can be set and removed. Tag filters are also available in the filter bar menu with all other applicable filters.

Clicking the “Tags” quick filter button or entering the “Tags” filter submenu from the filter bar will open the same menu with identical functionality. Clicking on a tag in this menu will add a “Tags” filter group with the specific item applied. This will cause only items with this tag to be shown in the interface (eg. if you add the EMPIAR tag as a filter on the projects view, only projects with the EMPIAR tag applied to them will now show in the card grid). Adding additional tags is an additive filtering process and will show all items with any of the added tags applied to them.

<figure><img src="/files/kTPDS9CBQx3CwJjzDlLB" alt=""><figcaption></figcaption></figure>

## Tag Management

<figure><img src="/files/Odv6Y4zRJ4WLJPG9kjOQ" alt=""><figcaption></figcaption></figure>

Tags can be managed in detail from the **Tags** tab within the management dialog. This view presents all tags in a table, displaying their relevant metadata, including ID, title, description, type, author, and creation date. Each row represents a single tag and includes action controls at the end of the row for editing or deleting that tag.

#### Filters

Several filtering options are available to simplify tag management. These controls are located in the top bar above the table.

* **Tag Type**: Filter the table to display only tags of a specific category (e.g., General, Projects, or Jobs).
* **Mine / All Toggle**: Limit the view to tags created by the current user or display all tags available in the instance.
* **Search**: Filter tags by direct text match. The search respects and combines with any other active filters.

#### Editing

Selecting **Edit** from the tag row’s context menu (accessible via the three-dot icon) opens the edit modal. From this modal, you can modify the tag’s title, update its description, and select a new colour.

<figure><img src="/files/95wsTGZEv5xXYvfqflPQ" alt=""><figcaption></figcaption></figure>

#### **Deleting**

Deleting a tag permanently removes it from the instance and from all items to which it has been applied. **This action is irreversible.**

#### Multi-Actions

The tag management table includes checkboxes in the first column, allowing multiple tags to be selected at once. Multi-selection respects all active filters and will only include tags currently visible in the filtered view.

Once one or more tags are selected, a **Delete {X} Tags** button appears in the top bar. Selecting this option permanently deletes all selected tags from the instance and removes them from any associated items. As with single-tag deletion, this action cannot be undone.

## **Tag Use Cases**

* Progression of a project by assigning various lifecycles (e.g., 'to-do', 'in-progress', 'done')
* Denoting the type of microscope used to collect the data present in a project (e.g., 300KV, Krios)
* Adding a demarcation of a quality result that can be referenced or returned to in the future (e.g., Good result)
* Using a tag for reconstructions of a specific particle type that can be referenced together across the instance regardless of project (e.g., `CoV S Protein` tag)


# Flat vs Hierarchical Navigation

Traditionally CryoSPARC has relied exclusively on a hierarchical navigation system, solely permitting workspaces to be viewed inside of their containing projects, and jobs to be viewed inside their containing workspaces. This has been reimagined in CryoSPARC v4 to allow navigating to each level (or granularity) of projects, workspaces, sessions, and jobs independently without having an upper level selection set. For instance, you could navigate directly to the jobs tab without selecting a project or workspace in order to see all jobs across the instance. This is nice to get a higher level view of what jobs are running, or have been run, in any shared projects or just generally to see the composition of your workflows. From here, you may want to open the filter options and filter all jobs available by type to compare results across the instance. By doing this, jobs can be aggregated and compared using a wide range of filtering options, either by their containers, or by their attributes.

Flattened views can be directly accessed using the “Spotlight” menu (accessible by clicking the magnifying glass icon in the navigation bar, the search bar button on the home page header, or by pressing the `command` + `k` keys together on the keyboard. From here you can simply begin typing “all” and options for viewing all projects, workspace, sessions, or jobs will appear below. Navigating to one of these options and pressing the `enter` key, or simply clicking on one, will navigate you to that granularity with no parent granularity selected (eg. all workspaces with no project selected, or all jobs with no project or workspace selected.

Flattened views are also automatically used for certain linking functions throughout the app to aggregate data with relevant filters applied automatically. An example of this is the “Job History” page. This can be accessed by clicking the overflow menu button with three horizontal dots in the navigation bar. The containing menu has an item called Job History in the first section with a clock icon beside the title. Clicking on the item will navigate you to the jobs granularity table view with no project or workspace selected and a status filter with statuses of completed, failed, and killed automatically applied. This was an entirely separate page in previous versions of CryoSPARC, but can now be built simply by adding the appropriate filters to the flattened jobs browse page. Further filters can be applied to this view if you would like to narrow the results further (such as a lane filter, user filter, and/or job type filter). This new aggregated data view can also be saved by bookmarking it in your browser or copying and saving the link constructed in the address bar. A CSV of the results shown in the filtered view can also be downloaded using the “Download CSV” button in the footer for archival or sharing purposes.

<figure><img src="/files/7z1t0PzSrihqZiGFToDo" alt=""><figcaption><p>The Job History view in the flattened jobs view with applicable filters applied.</p></figcaption></figure>


# File Browser

When selecting a path to create a project or input for a job parameter, CryoSPARC will display an integrated file browser.

The file browser supports [Python's glob syntax](https://docs.python.org/3/library/glob.html) for Unix style pathname pattern expansion.

<figure><img src="/files/kpfN4XgWvb1LBvdOBqXo" alt=""><figcaption></figcaption></figure>

## Search and Bookmarks

In v4.5+, the file browser allows for bookmarking directories for common access. To create a bookmark, right click on a directory and select 'Bookmark path' from the context menu that displays. Alternatively, click the bookmark button within the control bar to bookmark the current directory. Each bookmark can have a unique title, description and colour. Bookmarks are scoped to a specific CryoSPARC user account.

You can search and navigate to bookmarked directories and project directories from the sidebar.

<figure><img src="/files/qL1VFiwHsj2DijXqRNEH" alt=""><figcaption></figcaption></figure>


# Blueprints

## Introduction

Blueprints allow you to customize parameter settings of individual job types in CryoSPARC, and then save those customizations as templates to be used later in any project or workspace. The saved customizations, called blueprints, become available right from the CryoSPARC job builder and quick actions menus and can be easily created, applied, edited, and exported/imported to another instance.

As an example, if movies are frequently imported from the same microscope at the same magnification, an `Import Movies` blueprint could be created for that microscope to automatically set the Cs, pixel size, number of frames, electron dose, etc.

At a high level, blueprints are meant to create an organizational space that can be accessed quickly in any project and workspace to deploy relevant jobs from a curated template. They remove the need to maintain workspaces full of separate template jobs and alleviate the act of searching across the instance for a job previously used to process similar data that yielded good results; these jobs can be scattered across projects and workspaces and searching for them can be difficult and time consuming. Instead, with blueprints, templates of jobs are only a click away in the job builder.

## Creating a Blueprint

A blueprint can be created either from a pre-existing job, or from a job type (eg. `Import Movies`, `3D Classification`, etc). When creating a blueprint from a pre-existing job the Create Blueprint dialog will have it’s custom parameters pre-populated with the custom parameters of that job. When creating from a job type, the dialog will have no pre-set custom parameters.

### Creating from a Pre-existing Job

There are two ways to create a blueprint based on an existing job. First, by clicking the “Create Blueprint” option in the job’s quick access menu. Second, by selecting the job and clicking the “Create Blueprint” option in the job sidebar’s actions panel. Either of these options will open the Create Blueprint dialog. Similarly, with the job card selected, you can navigate to the job sidebar and open the actions panel by clicking the “Actions” button on the footer. You can select the “Create Blueprint” option in the panel to open the corresponding dialog.

From here you can add a title - which will be the name of the blueprint shown in the job builder - and modify or add any parameters you would like to have applied when using your blueprint. Clicking the green “Create” button at the bottom of the dialog will create your new blueprint.

<figure><img src="/files/PCRWLZlvTlUoW4g8Bf04" alt=""><figcaption></figcaption></figure>

### Creating from the Builder Sidebar Panel

You can also create a blueprint from scratch using the Job Builder. This will open the same creation dialog but with no custom parameters pre-populated. To do this, navigate to the Job Builder sidebar panel, find the job type from which you would like to create a blueprint, and use the “triple dot” overflow menu button to open the overflow menu; from here you can select the “Create Blueprint” option which will open the Create Dialog enabling you to build your blueprint.

<figure><img src="/files/dHGbqit2F20DjSdZ9ri4" alt=""><figcaption></figcaption></figure>

## Applying a Blueprint

To apply a blueprint, simply navigate to any relevant workspace and either choose a blueprint you would like to use from the Job Builder, or select a currently building job and apply the blueprint to it.

### Applying to a Pre-existing Job

Select the job you would like to apply the blueprint to, and using either the quick access menu or the job sidebar actions panel, select the “Apply Blueprint” option to open a submenu with relevant blueprint options for that job type. Selecting a blueprint will clear all of the pre-existing custom parameters (if any have been set) and apply all of the blueprint parameters.

<figure><img src="/files/6gXP6CCiaYCehe3Skb6Y" alt=""><figcaption></figcaption></figure>

### Applying from the Job Builder

Navigate into a workspace and then select the Job Builder sidebar panel. Find the job type that you would like to apply a blueprint from and open the blueprint drawer. The drawer can be opened by clicking on the down arrow beside the job type, or by navigating to it using the keyboard and pressing `shift` + `down arrow` to open the drawer. Select the blueprint you would like to apply. This will create a new job of the relevant type, and apply all of the blueprint parameters to it. From here it can have its inputs connected and be run as expected.

<figure><img src="/files/bvrYj9sgVVLzDuxxR8vm" alt=""><figcaption></figcaption></figure>

## Modifying a Blueprint

{% hint style="warning" %}
Blueprints cannot be edited or deleted unless you are the creator of the blueprint or an administrator.
{% endhint %}

Modification options can be accessed in the Job Builder by navigating to a blueprint and clicking the “triple dot” button on the far right side to open the overflow menu. This will present options to edit or delete the blueprint.

### Editing a Blueprint

Selecting the “Edit” option from the overflow menu will open an Edit Blueprint dialog which, for all intents and purposes, is the same as the “Create Blueprint” dialog. Any and all parameters can be added, reset, and removed, and the blueprint will be permanently updated with the new parameters by pressing the “Update” button at the bottom of the dialog.

### Deleting a Blueprint

Selecting the “delete” option from the overflow menu will open a confirmation popover which, upon confirming the deletion, will irreversibly remove the blueprint from the instance.

## Annotating a Blueprint

Blueprints are meant to facilitate consistency and repeatability in creating jobs for a specific purpose. This often means that a Blueprint will have a variety of specific parameters set to particular values. In order to maintain intelligibility over the reasons for setting those parameter values, and even the parameters themselves, Blueprints includes a suite of annotation options.

### **Job Details**

Job details operate nearly identically to those found in the job builder. They allow you to apply annotations on a job by job basis by adding a title and/or description. When creating a Blueprint these fields will be pre-populated with any pre-existing titles or descriptions that the original job was given. An additional option to the far right of the field label allows you to check or uncheck an “Apply to Job” option. This option determines whether the Blueprint title and/or description will be added to jobs that the Blueprint is applied to. These options are checked by default.

### **Parameter Details**

Each parameter includes an option for annotation. This can be accessed by clicking the “Additional Options” toggle button to the far right of the parameter.

<figure><img src="/files/54pX5iyyF1reswWMk71A" alt=""><figcaption></figcaption></figure>

From here you can add any relevant notes about the parameter, for instance why it was set to the current value and/or in what situations it should be changed. The parameter note will be visible inside of the sidebar information tooltip for the relevant Blueprint when hovered.

<figure><img src="/files/cY0rJqxdZaPoQJnUzbnG" alt=""><figcaption></figcaption></figure>

## Importing / Exporting a Blueprint

Blueprints, like workflows, are designed for portability. This means that they can be easily saved, stored, and shared between instances. Blueprints contain no identifying information of the instance that they were created in, and no references to the jobs that were used to create them. This allows blueprints to be a powerful tool for maintaining a catalog of your proprietary job templates, or cooperatively iterating on a processing approach agnostic of institution or instance.

### Exporting a Blueprint

A blueprint can be exported by clicking on the the “triple dot” overflow button beside the individual blueprint you would like to export. Clicking on the “Export” option in the menu will download the blueprint to your device. The downloaded file is a `.json` (JavaScript Object Notation) file and can be easily inspected using any modern web browser. This compact file includes all of the necessary information to recreate a blueprint in any CryoSPARC instance (running on a version current to or greater than the introduction of the blueprints feature).

### Importing a Blueprint

Importing a blueprint can be done by clicking the “Import Blueprint” button on the footer of the Job Builder sidebar panel. This will open a device native file browser where you can find and upload a previously exported blueprint `.json` file.

Once selected the file will be imported into your instance and will appear in the Job Builder sidebar panel like any other blueprint. The imported blueprint has no special properties outside of a `imported` attribute to demarcate it as created outside of the instance. The imported blueprint can be used, modified, and exported like any other blueprint.

## Summary

Blueprints are a template library for single jobs that allow you to consolidate useful template jobs in a single place where they can be easily created, applied, edited, and exported/imported agnostic of instance. They can be created from pre-existing jobs or from scratch using the Job Builder, and can be applied to pre-existing jobs or new ones. Blueprints can be modified by the creator or administrators, and can be exported and imported as `.json` files for portability.


# Workflows

{% embed url="<https://www.youtube.com/watch?v=Iz7V98aICIo>" %}

## Introduction

Workflows are a new system in CryoSPARC for quickly and easily populating a workspace with pre-defined sets of jobs. The system is designed with flexibility at its core. This means that it can be used to construct a top to bottom automated pipeline of jobs that can take you from import to refinement, a predefined branch of jobs for exploratory processing, or an arbitrary set of disconnected jobs and branches that can then be run independently or connected using preexisting systems.

Workflows can be self contained trees extending from import jobs, or branches dependent on any number of parent jobs created independently. This design allows workflows to build repeatable pipelines when the optimal path is known, while also retaining the ability to rapidly create branches for exploratory processing.

At a high level, workflows are designed to facilitate consistency when processing data using repeatable strategies. They are meant as a way to preload a series of jobs with their inputs connected and parameters set.

The process of using workflows is broken down into two steps. First a user must select a set of jobs as a template and use those to create the workflow. Once created the workflow can then be implemented in any number of project workspaces.

<div><figure><img src="/files/WIwizACwNXaYDKv69ki3" alt=""><figcaption><p>Creating a workflow</p></figcaption></figure> <figure><img src="/files/Zpx2HKI87cAqbzQFIfcm" alt=""><figcaption><p>Tree view containing three applied workflows</p></figcaption></figure></div>

## Creating a Workflow

We will use the [EMPIAR-10025](https://www.ebi.ac.uk/pdbe/emdb/empiar/entry/10025/) T20S Proteasome dataset and the general processing strategy presented in the [CryoSPARC introductory tutorial](https://guide.cryosparc.com/processing-data/get-started-with-cryosparc-introductory-tutorial) to demonstrate how workflows can be created and applied.

We will begin in a workspace with a successful run-through of this processing pipeline. If you have processed this dataset before, the job chain will likely look familiar. In this chain, we perform motion correction and CTF estimation of the movies, pick particles, clean with 2D classification, and finally perform ab-initio and homogeneous reconstruction.

<figure><img src="/files/zvW4VvNzQwXhyrlO2iQA" alt=""><figcaption></figcaption></figure>

First, we will select all of the jobs in this chain using either the [multi-select mechanic](https://guide.cryosparc.com/application-guide-v4.0+/managing-jobs#multi-actions) (`command/control` + `click`), or simply selecting the first and last job in the chain and using the “Select Job Chain” quick action.

<figure><img src="/files/6eNa5MCLD3DiAz2qJA39" alt=""><figcaption></figcaption></figure>

Once all of the jobs in the chain have been selected, the “Create Workflow” dialog can be opened by either clicking on the option in the quick actions menu or on the footer button in the multi-selection sidebar.

<figure><img src="/files/lchUmAItV2AgDax1eq39" alt=""><figcaption></figcaption></figure>

Using either of these options will open the “Create Workflow” dialog which is automatically populated with the selected jobs.

The “Create Workflow” dialog is broken into two panels:

<figure><img src="/files/STvMdPmNYeQ3FdwUPk7t" alt=""><figcaption></figcaption></figure>

* **Configuration Panel:**
  * A settings section at the top allows you to set a title, category, and description for the workflow.
  * Below the settings section is a sequential set of job panels with configurable details and parameters.

    <div align="right" data-full-width="false"><figure><img src="/files/BBFtzt9NjNxykUDNVK3v" alt=""><figcaption><p>Workflow job parameter</p></figcaption></figure></div>

    * Each parameter can be customized with a predefined value that is the default when the workflow is applied.
    * Reseting a parameter in this view will set it back to the job’s default parameter value.
    * Visibility can be toggled for each parameter, which defines whether or not the parameter can be seen when applying the workflow.
    * Locked status can be toggled for each parameter, which defines whether or not the parameter value can be changed from the default when applying the workflow.
* **Tree View**
  * The tree view is a graph representing all of the jobs and input/output connections in the workflow. It allows the workflow to be visualized more easily from a spacial perspective and can be used as a map for navigating the workflow jobs when setting parameters.
  * Job nodes are colour coded for legibility.
    * Default job nodes are set to a light grey with a blue accent for their ID.
    * The currently selected (building) job will be coloured purple.
    * Jobs that have modified parameters will be coloured green.
    * Disabled job nodes will appear light grey (these are nodes with no editable parameters).
    * Parent job nodes will appear light purple with a dotted border during creation of the workflow. During application of the workflow these jobs will be colour coded to their status (as denoted at the top of the configuration panel).
  * Clicking on a job node will select the job and automatically scroll the configuration panel to it.

To create our example T20S workflow we will add a title of “Processing Pipeline”, a category of “T20S”, and we will leave all of the default parameters as they are. To complete the creation process we will click the green “Create” button at the bottom right of the dialog.

## Applying a Workflow

Now that we’ve created our T20S workflow we can apply it. Navigate to the “Workflows” sidebar panel:

<figure><img src="/files/bfNkRk1ZF5zrVpv0LHO4" alt=""><figcaption></figcaption></figure>

From here we can click on our template to open the Apply Dialog:

<figure><img src="/files/zIQFmJxhLDTjjnXtYXTd" alt=""><figcaption></figcaption></figure>

The Apply Dialog is largely identical to the Create Dialog in composition and layout. The major differences reside in the left-hand configuration panel.

* The settings section contains a `Queue on Apply` option. This allows you to set all jobs to queue as soon as the template is applied. The `Queue to Lane` option allows you to choose the lane the jobs will queue onto during application if you toggled the `Queue on Apply` option.
* The proceeding job panels are structured very similarly to the Create Dialog and include all of the parameters that were exposed during the workflow’s creation.
  * Jobs that had no parameters set to a custom value or made visible during creation will not be shown, and will be coloured grey in the tree view.
  * Locked parameters are read-only in this view, and are demarcated with a lock icon.
  * Reseting a parameter in this view will set it back to the custom value defined during creation.
* The footer includes an “Apply” button to deploy the workflow into the current workspace, and a “Repeat” Button, which allows you to deploy the workflow and then open an identically configured Apply Dialog for quickly creating divergent branches or multiple exploratory pipelines.

For this example we will maintain our default parameter values, select the “Queue on Apply” option, and apply the workflow in the tree view by clicking the “Apply” button. All of the jobs included in the workflow will now be automatically created, connected together, and queued onto the selected lane.

<figure><img src="/files/UD258zONJkOdyU9pa0Vm" alt=""><figcaption></figcaption></figure>

## Modifying a Workflow

{% hint style="warning" %}
Workflows cannot be edited or deleted unless you are the creator of the workflow or an administrator.
{% endhint %}

Workflows can be modified from the Workflows sidebar panel by clicking the triple dot overflow button beside the individual workflow you would like to modify. This will open an overflow menu with multiple options for modifying and interacting with the workflow.

<figure><img src="/files/25m91J2jXY3QTdczGjkq" alt=""><figcaption></figcaption></figure>

### Editing a Workflow

Clicking the “Edit” option in the overflow menu will open the Edit Dialog. This dialog is identical to the Create Dialog and allows you to modify all parameters and settings to your liking.

Clicking “Save” will overwrite the Workflow with any changes you have made. Clicking “Save New” will save a new Workflow that is identical to the old one but with all of the changes you have made; you must change the Workflow title in order to use the “Save New” option.

### Deleting a Workflow

Clicking the “delete” option will open a confirmation popover which, upon confirming the deletion, will irreversibly remove the workflow from the instance.

## Rebuilding a Workflow

{% hint style="info" %}
If you are simply updating a Workflow’s parameters and/or annotations, it is recommended to use Edit functionality. Rebuilding is specifically designed to be used when fundamentally modifying input/output connections or adding/removing jobs from a Workflow.
{% endhint %}

Rebuilding a Workflow allows you to add or remove jobs from the Workflow, or change the input/output connections between jobs in the Workflow. If these jobs were created using the workflow that you are rebuilding, then all of their annotations will be carried over (notes, locked status, visibility status, etc.).

The best practice for rebuilding an existing workflow is as follows:

1. Apply a workflow into a new (or existing) workspace without queueing it.
2. Add new jobs, remove existing jobs, or change input connections as needed.
3. Select all of the relevant jobs that you would like to include in the rebuilt workflow.
4. Navigate to the workflow in the sidebar panel and select the “Rebuild” option from the triple dot overflow menu. This will launch the Rebuild Dialog, allowing you to see which jobs have been maintained or updated, and to create your rebuilt Workflow.

This system allows you to leverage the powerful tools that already exist in CryoSPARC for creating, deleting, and connecting or disconnecting jobs, while still retaining all annotations from the previous Workflow.

### Rebuilding Walkthrough

Enter a new or existing workspace in which you would like to edit your Workflow. Navigate to your Workflow using the “Workflows” sidebar, open the Apply Dialog, and apply the workflow into the workspace.

<figure><img src="/files/d3QX0arOXYkg2NgL7XCi" alt=""><figcaption></figcaption></figure>

Once the workflow has been successfully applied, you can modify it as you would any other set of jobs in CryoSPARC.

<figure><img src="/files/j7ItLPb49kPm3OI7YStX" alt=""><figcaption></figcaption></figure>

After modification is complete, you can select all jobs using any of the [multi-selection methods available](https://guide.cryosparc.com/application-guide-v4.0+/managing-jobs#multi-actions) and access the Create Workflow option through the multi-select sidebar footer or the job card quick actions menu.

<figure><img src="/files/pK0uI4heHmJfL4s8LOUi" alt=""><figcaption></figcaption></figure>

The Rebuild Dialog is largely identical to the Create Dialog, but includes a set of indicators at the top right of dialog in the header for more contextual information regarding differences from the original Workflow. Rebuilt jobs are those that existed in the original Workflow and have had all of their annotations copied over to the new Workflow. The “Rebuilt” and “New” indicators can be clicked on to open up a context menu with a list of all of the jobs pertaining to those statuses. Clicking on a job in the menu will select it and navigate you to it in the Workflow tree and sidebar.

<figure><img src="/files/zef8gY8K8fprojbg5g0P" alt=""><figcaption></figcaption></figure>

A Rebuilt Workflow must be saved as a completely new Workflow, distinct from the original. Rebuilding is treated as a fundamental modification of the original Workflow and so overwriting it is disallowed. We recommend adding a modifier to the title to indicate that this is a new version of the original (eg. “T20S Base Processing \[v2]).

## Workflows with Parent Connections

Workflows are designed to support both self-contained pipelines as well as exploratory processing. So far in this example we have been focused primarily on the former, so now let’s turn our attention to the latter.

Exploratory processing requires branching off from a singular pipeline in order to try different processing strategies simultaneously. For example, you may want to try building multiple refinement jobs of the same or different types, with a variety of divergent parameters set, and run them in parallel with the intention of seeing which strategy yields the best results. In this scenario, you might create a workflow containing a preferred setup of these jobs, beginning with a job to generate the initial volume (Ab-Initio Reconstruction) and a spread of refinement jobs.

In order to accommodate this style of processing, the workflows feature includes a concept of “Parent Connections”. Parent jobs are jobs that a workflow does not contain, but are required externally by the workflow. The parent to the aforementioned example workflow would be a job that outputs particles, such as a Select 2D, which would be connected into the Ab-Initio. Lets take a look at how this is done in practice.

### Creating a Workflow with Parent Connections

We will create a new workflow that begins with a “Select 2D” job as the parent, an “Ab-Initio” job as the first child, and a variety of refinement jobs connected to its outputs.

<figure><img src="/files/oIJTkSVfUHRk8uqvTVx4" alt=""><figcaption></figcaption></figure>

First we select the “Ab-initio” job and each of the refinement jobs using the multi-select mechanic (or the “Select Descendant Jobs” option from the quick actions menu). From here we will open the Create Dialog using the “Create Workflow” option in the quick actions menu. We now see our selected jobs in the tree view, with a “Select 2D” parent job connected to the “Ab-initio”. This parent job is indicated by it’s divergent colouring and “Px” ID (as opposed to “Jx”).

<figure><img src="/files/Tn0XzJczMBJDcboAX8CF" alt=""><figcaption></figcaption></figure>

The “Select 2D” parent job is a dependency of this workflow, because the “Ab-intio” relies on its inputs to function. The parent job will not be created by the workflow when it is deployed, but must be selected in order to allow the workflow to connect to its outputs.

### Applying a Workflow with Parent Connections

<figure><img src="/files/mL3eS5ySBjacoLqDm7vS" alt=""><figcaption></figcaption></figure>

We will now take our “Refinement Fork” workflow and apply it to a partially completed version of the T20s pipeline we worked on previously, that is currently stopped after a “Select 2D” job.

We will now use that “Select 2D” job to connect the particles output to our workflow by selecting it and then opening the workflow apply dialog from the sidebar. Once the dialog is open we can see that the “Select 2D” job is indicated as connected by the green notice at the top of the configuration panel, as well as its corresponding green colouring in the tree view.

Applying the workflow will create and connect all of the workflow jobs as expected, as well as automatically connecting the outputs of the parent “Select 2D” job with the created “Ab-initio” job.

We now have our “Refinement Fork” workflow deployed and queued. It has been automatically connected to our “Select 2D” job and will use the particle output to populate its input requirements.

#### **Missing Parent Connections**

If a workflow requires a parent connection to populate all of its necessary inputs, but no suitable parent job has been selected, it will display a “Missing Parent Connections” notice and the parent node in the tree view will be coloured orange.

<figure><img src="/files/QcJocRLKj365ujNaPh81" alt=""><figcaption></figcaption></figure>

You must exit the apply dialog by clicking the “Cancel” button or pressing the `Escape` key and select a suitable parent job within your workspace in order to create the connection. Once you have selected a suitable job, open the dialog by selecting the workflow from the sidebar again. It will now display a “Connected Parents” notice and the parent job node in the tree view will be coloured green.

<figure><img src="/files/vE4pXt7yrnrCwBVzX8Vx" alt=""><figcaption></figcaption></figure>

#### **Rerouted Parent Connections**

In some circumstances you may want to use a particular workflow with a different parent job than that which it was initially created, for example connecting an “Import Movies” job instead of an “Import Micrographs” job. Workflows includes a system to reroute input connections to facilitate this option. Any replacement job must output all of the required inputs and their necessary lower level slots that are consumed by jobs across the workflow. If these requirements are met, the original parent job will be replaced by the new selected parent job and all of the inputs will be connected as expected.

This system is automatic and does not need any configuration. If a suitable replacement job is selected it will be substituted in the workflow automatically.

<figure><img src="/files/hv4HQ6vBxyF5pdHkYQmw" alt=""><figcaption></figcaption></figure>

## Annotating a Workflow

Workflows are meant to facilitate consistency in complexity, and as such will often be composed of a variety of interconnected jobs with pre-set parameters. In order to maintain intelligibility, especially over time and between users, workflows include a suite of annotation options.

<figure><img src="/files/DSCVFhHM9v1yPZAXZFCR" alt=""><figcaption></figcaption></figure>

### **Workflow Details**

This is the highest level of annotation, and is meant to be used to give details relevant to the entire workflow. This may include links to related documentation, academic papers, and/or a relevant dataset. It can also be helpful to add any notes here about how the workflow is expected to be used. Workflow details can be added or changed in the configuration section of the Create Dialog or Edit Dialog.

These notes are displayed in the workflow sidebar tooltip that appears when hovering a workflow, as well as in the configuration section in the Apply Dialog in a “Details” panel that is closed by default.

<figure><img src="/files/na1o2SIt5XElQAMEvNgp" alt=""><figcaption></figcaption></figure>

### **Job Details**

Job details operate nearly identically to those found in the job builder. They allow you to apply annotations on a job by job basis by adding a title and/or description. When creating a workflow these fields will be pre-populated with any preexisting titles or descriptions that the original jobs were given. When applying the workflow the titles and descriptions can be edited in the Apply Dialog, and will be added to the jobs that are created by the workflow.

### **Parameter Details**

Each parameter includes an option for annotation. This can be accessed by clicking the “Additional Options” toggle button to the far right of the parameter.

<figure><img src="/files/YCqttWPQbolfFavhkFzK" alt=""><figcaption></figcaption></figure>

From here you can add any relevant notes about the parameter, for instance why it was set to the current value and/or in what situations it should be changed. When viewing this workflow in the Apply Dialog an info icon will be shown beside the annotated parameter title. When mousing over the icon the added note will appear inside of an info tooltip.

<figure><img src="/files/f2UF9n9BBydO90h6EWr2" alt=""><figcaption></figcaption></figure>

### **Flagging Parameters**

It can be desirable to add a requirement to change certain job parameters in a workflow to facilitate its use with different datasets. There are many situations where a parameter will almost always need to be changed in order to run the workflow successfully, for example import paths or box size. Workflows allows you to flag a parameter as a pseudo-requirement to address this situation.

You can access the flag button from the same “Additional Options” parameter toggle button used to add parameter notes. By clicking the flag button in the top right of the panel you can flag the parameter.

<figure><img src="/files/XB00dRPGeFBtc0iW1q8j" alt=""><figcaption></figcaption></figure>

When viewing the Apply Dialog the flagged parameter will now have an orange outline and a flag icon beside its title, the corresponding job node in the tree view will be accompanied by a flag icon. On the dialog footer you will see a “Flagged Parameters” tracker that shows the total number of flagged parameters and number of updated flagged parameters. Clicking this button will reveal a menu checklist of flagged parameters, organized by job, with a check mark circle indicating whether it has been updated or not. Clicking a parameter in the menu will navigate you to it.

<figure><img src="/files/9Fe488Jnf3VPsP4uzryg" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
It is important to note that flagged parameters are not hard requirements. A workflow can be deployed without its flagged parameters having been updated. This is to maintain flexibility and not lock a user into updating a parameter they do not wish to.
{% endhint %}

Once you have added a title and modified included job parameters to your liking, you can navigate to the dialog footer and click the “Create” button to create the workflow.

## Importing / Exporting a Workflow

Workflows are designed from the ground up for portability. This means that they can be easily saved, stored, and shared between instances. Workflows contain no identifying information of the instance that they were created in, and no references to the jobs that were used to create them. This allows workflows to be a powerful tool for sharing successful data processing pipelines or strategies, maintaining a catalog of your proprietary processing workflows, or cooperatively iterating on a processing approach agnostic of institution or instance.

### Exporting a Workflow

A workflow can be exported by clicking on the the triple dot overflow button beside the individual workflow you would like to export. Clicking on the “Export” option in the menu will download the workflow to your device. The downloaded file is a `.json` (JavaScript Object Notation) file and can be easily inspected using any modern web browser. This compact file includes all of the necessary information to recreate a workflow in any CryoSPARC instance (running on a version current to or greater than the introduction of the workflows feature).

### Importing a Workflow

A workflow can be imported by clicking the “Import Workflow” button on the footer of the Workflows sidebar. This will open a file browser where you can find and upload a previously exported workflow `.json` file.

Once selected the file will be imported into your instance and will appear in the Workflows sidebar like any other workflow. The imported template has no special properties outside of a `imported` attribute to demarcate it as created outside of the instance. The imported workflow can be used, modified, and exported like any other.

## Pinning

Any Workflow that has a parent connection can be pinned to the quick actions menu, allowing it to be accessed in the menu of any job matching the Workflow’s parent job type (eg. if the Workflow has a parent connection to a *Patch CTF Estimation* job, pinning that Workflow would allow you to access it through the quick actions menu of any *Patch CTF Estimation* job).

In order to pin a Workflow, navigate to it in the Workflows sidebar and open its overflow menu from the triple dot button. Click the “Pin” option from the menu. The Workflow will now appear in the quick actions menu for any job with the parent job’s type. You can “Unpin” the job from the overflow menu the same way, if you no longer want it to appear in the quick actions menu.

{% hint style="info" %}
The pinned Workflow will only appear in the quick actions menu for exact matches of the parent job’s type. If you would like to[ ](https://guide.cryosparc.com/application-guide-v4.0+/workflows#rerouted-parent-connections)[reroute the Workflow’s parent connections](https://guide.cryosparc.com/application-guide-v4.0+/workflows#rerouted-parent-connections)[ ](https://guide.cryosparc.com/application-guide-v4.0+/workflows#rerouted-parent-connections)from a different job type, you will need to select the job and open the Apply Dialog from the sidebar.
{% endhint %}

Using the pinned quick action will still open the Apply Dialog, allowing you to set all preferred options before applying into a workspace.

## Summary

CryoSPARC's Workflows allow for the creation of pre-defined sets of jobs for data processing. Workflows can be self-contained trees or branches dependent on parent jobs. They facilitate consistency in processing data and can be used for pipelining known quantities or breaking up different steps during exploration. Workflows can be created by selecting jobs and opening the "Create Workflow" dialog, which allows for customization of parameters and visibility. Workflows can be modified and deleted, and can be exported and imported as JSON files.

## Limitations

Jobs that create their outputs while in `running` status rather than when in `building` status cannot be automatically linked together when applying a Workflow. This will cause any job chain to break where these jobs are present. This is a fundamental limitation of Workflows as they must create all input/output connections on `building` jobs.

Current jobs known to exhibit this behaviour are:

* **Import Result Groups:** This job type cannot be used directly in a Workflow, however it may still be used as a parent connection.
* **Flex Generate:** Can still be used in a Workflow without issue so long as it does not have any jobs connected to its outputs

You will be warned by the Workflow if these jobs are present as parents, and you will be unable to create the Workflow until they are updated or removed from the selection.

There are also jobs that will exhibit this behaviour if they were created in earlier versions of CryoSPARC:

* **Import Particle Sets:** Includes a parameter `Enable Strict Checking` which was added in v4.4.0 and needs to be toggled on in order for outputs to be generated when `building`. This parameter is on by default in newer CryoSPARC versions.
* **Exposure Group Utilities:** Has been updated in v4.4.1 to generate outputs while in `building` status instead of in `running` status. This will not be an issue for instances running versions after v4.4.0.

Jobs which create their outputs while in `running` status will be disconnected from subsequent jobs created by the Workflow. To create a workflow which includes this type of job, you can create separate workflows for before and after these jobs. The first Workflow would include all jobs up to the job that generates its outputs while `running`, and the second Workflow would use that job as a [parent connection](https://guide.cryosparc.com/application-guide-v4.0+/workflows#workflows-with-parent-connections) and run the rest of the jobs in the pipeline. In this scenario you would apply the first Workflow and allow all jobs to run to completion, and then select the final job and apply the second Workflow.

In most cases it is *possible* to simply run a single workflow until the job in question is complete and then connect its generated outputs to the disconnected branch of the Workflow. However, re-connecting the workflow in this way can cause certain passthroughs to be missing in jobs downstream of the re-connected jobs. This can lead to downstream jobs failing due to not having access to required input group slots. It is therefore not the recommended way of handling jobs which create their outputs in the `running` status.


# Managing Jobs

## Job Management

### Current Jobs Dialog

<figure><img src="/files/z0IWDEGDJ8afbA6IuWBw" alt=""><figcaption><p>In v5.0+, lanes and individual jobs display up-to-date resource allocation data</p></figcaption></figure>

This is the hub for managing jobs and viewing instance usage. The Current Jobs panel shows all jobs in progress in a grid of their containing lanes.

The control bar at the top of the view contains a view toggle for filtering between all jobs, queued jobs, or active jobs. On the right side of the bar, instance token usage is visible (if applicable). Token usage is broken down by each active instance and summarized in total.

### Lanes

The main area is composed of lanes and their attached jobs. This view has been condensed from earlier versions of CryoSPARC by moving all inactive lanes into a single tracker pinned to the left side of the viewing area. Only lanes that are currently active will be displayed separately. Each lane is composed of a three parts.

<figure><img src="/files/s3RvyqNxJaqZrk9zcJWh" alt=""><figcaption></figcaption></figure>

**Header**

The header shows the lane title, type, and the number of jobs currently attached to the lane. The job number will change if viewing jobs filtered by active or queued, and will show what number out of the total are currently being displayed.

The header also houses additional information visible when clicked to expand. This will reveal the title, name, and description of that lane.

In v5.0+, the the lane header will display the total sum compute resources available on that lane and how much resources are currently used by jobs:

<figure><img src="/files/VogTZojvWI1N6VQyRkjo" alt="" width="299"><figcaption></figcaption></figure>

**Jobs**

<figure><img src="/files/T5pICPguYicHeq2fpTrL" alt="" width="383"><figcaption></figcaption></figure>

The jobs section houses unique job cards with features specific to managing the instance.

The card header shows the status colour indicator followed by the job’s project ID and individual job ID. On the right side is the number of GPU tokens being used by the job, its running priority, and its current status. The job type is displayed below.

The job body includes information about who created it and how long it has been attached to the lane in its current status. Below is an action button to kill the job, and a trigger button for the job’s quick access menu.

The job card can also be selected and inspected in the sidebar, and opened from the “View Job” button at the bottom of the sidebar (to be noted: this will close the manage dialog as only one dialog can be open at a time, but it is possible to go back to the manage dialog by simply pressing the native browser back button).

In v5.0+, the job card includes compute resources allocated in the footer. It can be viewed by clicking the header:

<div align="center"><img src="/files/Jq9TvYigTadaPH3b7DZ4" alt=""></div>

**Footer**

The lane footer contains a “View Details” button that will navigate to the flattened jobs browse section set to the table view, and will set filters for the relevant lane and job statuses of “Queued, Launched, Started, Waiting, and Running”. This allows for more detailed and dynamic inspection of jobs running on the lane.

### **Main App Footer**

<figure><img src="/files/qAKpyFD90leCVZ9no350" alt=""><figcaption></figcaption></figure>

The app footer is designed to give contextual information about the jobs currently active in the instance at a glance. It is highly dynamic and will size up or down depending on the space available in the window. The footer is made up of two sections, the active job section and the active target section.

* **Active Jobs:** The active job section is composed of a details button which is followed by job pill buttons. The details button shows a count of how many jobs are currently queued and/or active. This button will open up the “Current Jobs” dialog when clicked. The job pill buttons show the job’s project ID, job ID, and a status indicator. Hovering a job button will show the full job type in a tooltip, and clicking on a job will open that job’s inspection dialog (this will not navigate you away from your current page).
* **Active Targets:** The active targets section has its own details button which similarly opens the “Current Jobs” dialog. The target buttons show the applicable lane’s name and can be clicked to open a target dialog with a small grid of job cards representing the jobs running on that target. Each card has nearly identical functionality to those in the “Current Jobs” dialog.

## **Job Actions**

All jobs have variety of actions available for both processing data as well as doing general management and maintenance. These actions can be accessed from two distinct locations, the quick actions menu, and the sidebar actions panel.

**Sidebar Actions Panel**

<figure><img src="/files/OBVK2IKImQzWYO14XmqX" alt=""><figcaption></figcaption></figure>

The sidebar actions panel is analogous to the actions section in the jobs sidebar of previous versions of CryoSPARC. This panel can be expanded by clicking the “Actions” button on the footer of the job sidebar. The actions panel overlays the main sidebar area, and allows you to continue to see and scroll through all of the information available in the sidebar. This panel is focused on managing the current job’s data and only includes core job actions (these are detailed below). This panel will remain open until you decide to close it, which allows for ease of iteration while processing data within a job whether the job dialog is open or closed.

**Quick Actions Menu**

<figure><img src="/files/QZnJkk9m54XciMczBsIc" alt=""><figcaption></figcaption></figure>

The quick actions menu is a context menu that can be opened on job cards or rows and gives a wide range general actions and job specific actions. In order to access this menu you must either right click on a job card or row, or click the triple dot menu button on the job card header. The quick actions menu has five sections: the core actions section, multi actions section, quick actions section, general actions section, and navigation section. [The quick actions section is populated by common next steps in a processing chain, and allows you to quickly build subsequent jobs with the inputs already connected. ](/application-guide/creating-and-running-jobs#job-quick-actions)[The general actions section includes the tags menu where you can add or remove tags](/application-guide/tags)<mark style="color:red;">.</mark> The navigation section has options to enter the job’s containing project and view all workspaces inside, or enter the job’s workspace and see all jobs inside. The core actions and multi actions sections are covered below.

### Core Actions

#### **Queue Job**

Queuing jobs will connect them to a selected lane and set their status to queued until the necessary resources become available to run them. This is covered further in the “[Creating and Running Jobs](/application-guide/creating-and-running-jobs)” guide section.

#### **Link Job**

To avoid redundant processing, it's possible to link jobs across multiple workspaces to continue processing in a new or existing workspace. Any changes you make to the job in one workspace will be reflected in the other, as both workspaces will be displaying the exact same job. If a linked job is deleted from any workspace, all other instances of it in other workspaces will also be deleted. You can unlink a job from a workspace if needed.

#### **Move Job**

When re-organizing your workspaces, you may find it necessary to move a job from one workspace to another. This action will cause the job to disappear from its original workspace and move to the new one.

#### **Kill Job**

You can kill a job while it is running. A pop-up message will allow you to confirm the action. Killed jobs will display as orange and will remain as part of the workspace until you delete them. Killed jobs will not have their intermediate results expunged until they are cleared.

#### Restart Job

Any job that is in a killable or clearable state can be restarted. This action will run through a process of automatically killing the job, clearing the job, and then queueing the job to the same lane with the same parameters it was initially run with.

#### **Clear Job**

When a job has been killed, has failed, or has been completed, it can be cleared. This means that its results and outputs are all erased from the file system, but its connections to other jobs and parameters are retained. This makes it easy to correct a mistakenly set parameter, or to re-run a failed job. Clearing also removes queued jobs from the queue and resets jobs to Building status.

#### **Clear Intermediate Results**

Unused intermediate results can be cleared from iterative jobs to save space. This action can be executed at a job level (by clicking "Clear Intermediate Results" in the quick actions menu or actions panel) or at a project level for every job (by clicking "Clear Intermediate Results" in the project level quick actions menu or actions panel). This function will remove all unused outputs created by iterative jobs that save raw data at every iteration. Final results for every result slot will be retained, whether they have been used elsewhere or not.

#### **Export Job**

Any individual job can be exported, for sharing, manipulation, or archiving. Jobs must be exported manually in order to create a "consolidated" exported-job directory which is then importable. For more details, please see: [Exporting a job](https://guide.cryosparc.com/application-guide/pages/-MNiDCV_2xPWJWCTbMK7#5.-ability-to-export-and-import-individual-jobs).

#### **Clone Job**

Cloning is particularly useful when you wish to process subsequent jobs in a new workspace for better organization of your experiment, or if you wish to quickly replicate a job (to try another setting of parameters) without dragging and dropping inputs into the Job Builder. Select the job you wish to clone, and from the quick action menu or sidebar actions panel, click 'Clone job'. By default, the job will be cloned in the current workspace. Click 'Queue' to launch the job. Note that you can always edit the parameters of a cloned job before Queuing. You can also change inputs by removing connected inputs and dragging in new inputs.

#### **Mark Job as Complete**

This option allows you to mark failed or killed jobs in CryoSPARC as completed in order to allow their latest outputs to be used for further processing. For more details, please see: [Data Management in CryoSPARC](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/tutorial-data-management-in-cryosparc#5-ability-to-export-and-import-individual-jobs).

#### **Delete Job**

While selected, you can delete a job from the quick actions menu or sidebar actions panel. A pop-up message will ask you to confirm the deletion. Once deleted, the job will disappear from the workspace (except in the tree view, if it has children that need to be displayed). Note that job IDs within a workspace are unique, so a new job in the same workspace will not be assigned the same ID as a previously run or previously deleted job. Deleted jobs are cleared before deletion, meaning that their intermediate and final results are erased from disk.

#### Mark Job as Final

Jobs can be marked as final for the purpose of retaining their data during certain data cleanup operations. This and related functionality is explained in depth in: [Guide: Data Cleanup (v4.3+)](https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/guide-data-cleanup-v4.3+).

### Multi Actions

Multi actions are actions that only appear when multiple jobs have been selected. This can be done by holding the `command` key and clicking on job cards or rows to select multiple, or by holding the `command` key and pressing or holding either the `left` or `right` arrow keys with one or more jobs already selected.

Once multiple jobs have been selected a new multi-actions menu will become available in place of the default actions menu when right clicking on job cards or rows. A multi-actions palette is also available and can be activated by clicking the “Actions” button at the bottom of the multi-select sidebar.

Jobs can be removed from the selection by either clicking the “clear” button in the top right corner of the job cards in the multi-select sidebar, or by holding the `command` key and clicking selected job cards or rows.

All selected jobs can be cleared by either clicking the “clear” button on the righthand side of the multi-select sidebar header, clicking the background of the cards view, or pressing the `escape` key.

<figure><img src="/files/3cF5lzywFXFheN4dOY0k" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/ZcHNsa08N8UO9hNuhZlL" alt=""><figcaption></figcaption></figure>

#### Queue Jobs

This action allows you to select multiple applicable jobs and choose a lane to queue them too. All selected jobs will be queued to the chosen lane, with the exception of interactive jobs and master direct import jobs, which will automatically be queued to the default lane even if a different lane was chosen.

#### Kill Jobs

This action will kill all applicable selected jobs, and works identically to the singular “Kill Job” action.

#### Clear Jobs

This action will clear all applicable selected jobs, and works identically to the singular “Clear Job” action.

#### Restart Jobs

This action allows you to restart any selection of jobs that are in a killable or clearable state. It will automatically kill any running jobs, clear all stopped jobs, and then queue each job to the same lane with the same parameters it was initially run with.

{% hint style="info" %}
Only a single restart process can be performed at a time in order to prevent jobs states from becoming inconsistent during the operation.
{% endhint %}

#### Mark Jobs as Complete

This action will mark all applicable selected jobs as complete, and works identically to the singular “Mark Jobs as Complete” action.

#### Link Jobs

This action allows you to link all selected jobs to another workspace. When selecting this option a confirmation alert will be shown with a list of all selected jobs for review as well as options to choose the workspace where the selected jobs will be linked, or create a new workspace to link the jobs to. When the link is confirmed the jobs will be linked to the chosen or new workspace and you will be navigated into the workspace with the linked jobs selected.

#### Move Jobs

This action functions similarly to the “Link Jobs” action and allows you to move all selected jobs into another workspace. When selecting this option a confirmation alert will be shown with a list of all selected jobs for review as well as options to choose the workspace where the selected jobs will be moved, or create a new workspace to move the jobs into. When the move is confirmed the jobs will be moved into the chosen or new workspace and you will be navigated into the workspace with the moved jobs selected.

#### **Clone Jobs**

This action allows you to clone a variable number of jobs based on how many you have selected. When this option is chosen an alert will open showing you the jobs to be cloned with an option to clone them into either the current workspace, another workspace, or a new workspace. When cloning is confirmed the jobs will be cloned into the chosen or new workspace, if the chosen workspace is different from the one you are currently in you will be navigated into it. The cloned jobs will be selected automatically after the action is complete.

{% hint style="info" %}
When selecting “Clone Jobs” the alert will open with the “Clone” confirmation button automatically focused. If you wish to clone the selected jobs into the current workspace, you can simply press the `enter` key to do this in an expedient manner.
{% endhint %}

#### **Clone Job Chain**

This action allows you to clone a chain of jobs between two selections. This can only be done with a valid chain of jobs, if unrelated jobs are selected as the start and end points you will be alerted that the chain is not valid. When this option is selected a dialog with the same options as the “Clone Jobs” action will open. Here you can see all of the jobs that would be created in the chain (this list is all of the jobs between the two selected points), and options to select a workspace to clone into or create a new workspace to clone into. When cloning is confirmed the job chain will be cloned into the chosen or new workspace, if the chosen workspace is different from the one you are currently in you will be navigated into it.

#### Clear Intermediate Results

This action will clear the intermediate results of all applicable selected jobs, and works identically to the singular “Clear Intermediate Results” action.

#### Export Jobs

This action will export all applicable selected jobs, and works identically to the singular “Export Job” action.

#### Unlink Jobs

This action will unlink all selected linked jobs from the current workspace, and works identically to the singular “Unlink Job” action.

#### **Delete Jobs**

This action allows you to multi-select any number of jobs and delete all of them. A confirmation alert will open once this option is selected with a list of all of the jobs to be deleted available for review. Once confirmed, this action will permanently delete all selected jobs.


# Interactive Jobs

## Overview

Interactive jobs represent the points in the CryoSPARC data processing pipeline where user interaction and intervention occur. These are either required interactions to build outputs necessary in sequential steps or optional interventions to operate on data manually in order to secure better results downstream.

Interactive jobs always run on the master node and when queued will run until reaching a stage of necessary intervention, which is represented by the “waiting” status. At this point the job will gain a new “Interactive” tab inside the job dialog. This tab will be the default view when opening an interactive job with the “waiting” status, and will be switched over to automatically if the dialog is already open when it becomes available.

{% hint style="info" %}
Interactive jobs have been designed to persist their state in the browser’s session storage. This means that all active interactive jobs will save any changes you make to their configuration, and will persist these changes even if you close the job or refresh your window. These changes are isolated to the current browser tab, and will be cleared if the tab is closed. Similarly, if you would like to look at job with the default setting applied, you can simply open it in a new tab or window.
{% endhint %}

All major interactive jobs share the same two column layout. The overview column is on the lefthand side and the inspection column is on the righthand side. Certain panels within these columns are foundational systems for interacting with these jobs and are present across interactive job types.

## **Exposure Plot**

The exposure plot is a generalized plot with two view options (scatter or histogram) that shows all of the exposures present in the job with configurable axes. This plot is designed to allow for quickly visualizing patterns and trends in your data that you can further act on with the other tools available in the job.

The plot type is configurable using the “type” selection menu in the top left of the panel above the plot. Both plot types have a download button located in the top right of the panel. This button allows you to download an image of the plot to share or keep for your records.

### **Scatter Plot**

<figure><img src="/files/XpUtZUaKqeRRHeQJixcj" alt=""><figcaption></figcaption></figure>

The scatter plot option shows all exposures across two configurable axes. Each point in the plot can be hovered to reveal an exposure preview that opens offset on the righthand side of the plot. These previews show a number of pertinent fields from the exposure and an exposure preview at the top if available. Each point in the scatterplot is also an actionable button, and will select the corresponding exposure, scroll to it in the micrographs table, and set it in the micrograph viewer when clicked.

{% hint style="info" %}
All exposures processed through the *Patch Motion Correction* job will have thumbnail previews available. To generate micrograph thumbnails for existing datasets, use the *Generate Micrograph Thumbnails* job.
{% endhint %}

### **Attribute Histogram**

All interactive jobs feature a configurable exposure plot with histogram option to view the distribution of a particular attribute. Exposure curation supports a 'split histogram' chart which displays distributions of accepted and rejected exposures as separate traces.

<figure><img src="/files/Ai4QBLck1cGayVqq6cjD" alt=""><figcaption></figcaption></figure>

## **Micrographs Table**

<figure><img src="/files/MA3QryITJtbADdiKYeoR" alt=""><figcaption></figcaption></figure>

The micrographs table shows all micrographs present in the job with a number of sortable columns for pertinent fields (eg. DF Average, Relative Ice Thickness, CTF Fit).

The table header has three actionable items. The scroll target button allows you to quickly scroll back to a selected table row in case you have scrolled the table to another location. The filter menu allows you to select subsets of micrographs based on status (eg. accepted, rejected, exposures with picks, exposures without picks). The final item is a CSV download button that allows you to download a CSV of the entire table data with any sorting or filter options applied.

## **Micrograph Viewer**

<figure><img src="/files/wOgKGlioOOI79qbsmdGp" alt=""><figcaption></figcaption></figure>

The micrograph viewer will be immediately familiar to anyone that has used CryoSPARC Live, and has nearly identical functionality. It consists of a viewing plane with a number of overlayed navigation and action palettes, as well as some data modules with contextual information.

The micrograph viewer control bar contains a variety of options for displaying and controlling the data visualization options on the micrograph. These include toggling particle diameter and box size on or off, as well as adjusting their respective sizes; also available are options for controlling the contrast intensity override and changing the pick or box colour. Advanced controls are contained in the triple dot overflow menu button.

<figure><img src="/files/yhqj1NidRmKYCW3zWLOW" alt=""><figcaption></figcaption></figure>

The micrograph viewer will initialize with the micrograph fit to the viewer window. Your cursor will initially be set to drag mode which will allow you to click and hold on the micrograph to drag it around the viewer. Scrolling in the viewer will zoom into the micrograph at the location of your cursor.

* **Scale Bar**: The top right corner of the viewer houses the scale bar. This bar shows a dynamic scale that changes as you zoom in or out, and will switch colours to maintain visibility over your micrograph.
* **Lowpass Filter:** This slider allows you to easily change the lowpass filter (in angstroms) to gain better contrast of the micrograph.
* **Download Button:** This button allows you to download an image of the micrograph with any overlaid information present. This includes picks and boxes set over particles.
* **Scale Palette:** The scale palette is located in the bottom left corner and allows you to quickly refit the micrograph to the viewer display area. It also gives multiple preset scale option buttons to quickly jump to a specific zoom level without scrolling.
* **Navigation Palette:** The navigation palette is located in the bottom right corner and is split into two modules. One controls scroll behaviour, and allows you to set scrolling to either zoom or pan the viewer. The other controls the click behaviour and allows you to switch between the measurement tool and dragging tool. The measurement tool allows you to click and hold on the micrograph to drag out a circular measurement ruler. This ruler will show a red circle with a tooltip indicating the circle’s diameter in both pixels and angstroms.
* **Action Palette:** This palette is only available in some jobs and allows actions to be taken on the micrograph for data processing purposes (eg. for the manual picker you can change between click options for adding and removing manual picks).

## **Interactive Job: Manual Picker**

<figure><img src="/files/lVuNxnPgm4yoDDIhWNrk" alt=""><figcaption></figcaption></figure>

The manual picker job allows you to pick particles interactively and extract particles for generating templates.

The control bar at the top of the job allows you to see a quick statistical overview of the running job, including the total number of exposures, total exposures with picks, total exposures without picks, and the ratio between the set particle diameter and box size. The “Done Picking | Extract Particles” button sits at the far right side of the control bar and will complete the job and extract all particles when clicked.

The main content area breaks down into two panels, the overview panel, and the inspection panel.

### Overview Panel

The manual picker overview panel contains the standard configurable scatterplot at the top and the micrograph table below. These are covered in more detail in the previous section.

### Inspection Panel

The inspection panel includes a micrograph viewer with various overlayed controls, and a control bar above the viewer that allows you to configure picking and viewing parameters.

The main bar includes stacked particle diameter and box size configuration inputs. Each of these allows you to toggle on or off the picks that appear on the exposure (this is a visual toggle only and will not remove picks), and an input and slider that allow you to adjust the size of the picks. The particle diameter is used solely for visualization purposes, while the box size is used for extraction. If the provided maximum value on the slider is not sufficient for your required selections, you can adjust the number in the input to any value and the slider will automatically update its bounds to accommodate that number.

The control bar also includes an actions button in its top right corner that will open a menu with additional parameters when clicked. Included are the “Contrast intensity override” and options to adjust colour of the picks or boxes that appear on the micrograph, as well as a toggle to hide or display the scale bar in the exposure viewer.

## **Interactive Job: Inspect Particle Picks**

<figure><img src="/files/Zwybvd7q5nPEmtVfGfYP" alt=""><figcaption></figcaption></figure>

The inspect particle picks job allows you to inspect and modify picks using various thresholds following auto-picking.

### Overview Panel

The inspect picks overview panel contains the standard configurable scatterplot at the top and the micrograph table at the bottom. These are covered in more detail in the previous section.

This job additionally contains a Power Histogram panel which includes an image of the power histogram with Power Score and NCC Score input sliders colocated along its relevant axes for clarity and ease of visualization. The filament histograms (Local Curvature and Filament Sinuosity) also include colocated input sliders along their relevant axes.

### Inspection Panel

Along with configuring a particle diameter to view picks, you can set a box size and choose to overlay it within the exposure viewer. This functionality is virtually identical to that of the manual picker, except that the selections are not used for any outputs, but solely for visualization purposes.

<figure><img src="/files/CpPo48N8NbLZtT5a6HTe" alt=""><figcaption></figcaption></figure>

### Auto Clustering

As of CryoSPARC v4.6+, Inspect Particle Picks has a mode to cluster particles based on their particle picking scores (NCC and power score) and automatically select a cluster that will ideally contain true particle picks. This feature allows Inspect Particle Picks to be run fully automatically, without user input.

The auto clustering mode is designed to work with particle picks that were picked using denoised micrographs from the [Job: Micrograph Denoiser (BETA)](/processing-data/all-job-types-in-cryosparc/exposure-curation/job-micrograph-denoiser-beta). It operates by first separating particles into five defocus bins spanning the defocus range of micrographs in the dataset. Within each bin, clustering is performed on the pick stats (NCC and power score) and the cluster with the average power score closest to the "Target power score" parameter (default 50) is selected as the cluster to keep. The selected clusters from the five defocus bins are combined to produce the final output.

<figure><img src="/files/61YUAYjCIfLg60rn2Ghx" alt="" width="563"><figcaption></figcaption></figure>

To enable the Auto-Clustering mode, turn on the Auto Cluster parameter. It may be necessary to adjust the target power score parameter for certain datasets depending on e.g. the particle size. The target power score settings should be re-usable across multiple datasets of the same or similar particle.

<div data-full-width="false"><figure><img src="/files/5NzrcvBrGyG8gutXHuk1" alt="" width="375"><figcaption></figcaption></figure></div>

## **Interactive Job: Manually Curate Exposures**

<figure><img src="/files/pUOSv3oeOMe0IKR75pmo" alt=""><figcaption></figcaption></figure>

Manually curate exposures enables a user to visually inspect and curate a set of micrographs or movies using either population-level statistics and adjustable thresholds, or individual inspection of diagnostic plots.

### When to use this job

The overall goal of this interactive job is to aid in cleaning up your data so that low-quality micrographs, and their associated particles do not detract from achieving the highest possible resolution refined structures.

**After pre-processing:** After `Patch Motion Correction` and `Patch CTF estimation`, you can use `Manually Curate Exposures` to help you identify and exclude micrographs with sub-optimal characteristics, e.g., broken or thick ice, carbon edges, too much motion, etc.

**After particle picking and `local motion correction`:** Once you have picked particles and estimated location motion trajectories, you can use this information in `Manually Curate Exposures` to identify and exclude micrographs associated with those picks whose local motion trajectories deviate from the norm for your dataset. In this case, the job will output a filtered set of micrographs as well as particles.

**After using the Junk Detector**: If micrographs have junk annotations, the labeled areas will be displayed and colored by junk type. Micrographs can be filtered by the percentage of their total area that is covered by junk of various categories.

<figure><img src="/files/mf0BQJ5PoUYUDPbXPBoi" alt=""><figcaption></figcaption></figure>

### **Job Overview**

The control bar at the top of the job gives a quick overview of relevant statistics on the lefthand side, and houses the “Done” button for completing the job on the righthand side.

The main panels are arranged similarly the Manual Picker and Inspect Picks jobs, with the overview panel on the lefthand side of the main content area, and the action panel on the righthand side.

### **Overview Panel**

The same functionality exists in the overview panel as in the aforementioned jobs, with a configurable scatterplot at the top and a micrographs table of all available exposures below (in this case thumbnail previews and metadata shown when hovering plot points are loaded from Motion Correction). Exposure Curation also supports a separate cards view that can be used to quickly examine and manually reject/accept individual micrographs.

#### Exposure Card View

The exposure card view is available by clicking the card view button on the control bar above the Micrographs Table. This will transition the table to a grid of exposure image thumbnails.

<figure><img src="/files/1ZrjNYPLyu4LCJ6beMk6" alt=""><figcaption></figcaption></figure>

The grid automatically fills each row with as many exposure cards as possible while maintaining aspect ratio and maximum resolution.

From the top left and clockwise, each exposure card shows the exposure’s index, defocus average in micrometers, the rejection status if applicable (an “M” represents manual rejection and a lack of one represents threshold rejection), and a reject/accept button visible on hover (clicking this button will toggle the exposure’s rejection status). The index and defocus average tags can be hidden using the “Details” toggle on the control bar.

Clicking on an exposure card will select it and switch the action panel to the Micrograph tab (revealing the Exposure Viewer). Once an exposure has been selected the keyboard can be used to navigate between exposures using the left and right arrow keys. Pressing the “R” key will toggle the exposure’s status between accepted and manually rejected. Hovering over an exposure card will reveal the exposure preview rendered directly beside the grid with a variety of pertinent information.

### **Action Panel**

The action panel is configured with a main action bar at the top, housing tabs for changing the panel view, as well as navigation arrows for moving between individual exposures, and a rejection button for manually rejecting a selected exposure.

<figure><img src="/files/yl6En4RZ86qY38zZiMpj" alt=""><figcaption></figcaption></figure>

“Thresholds” is the default panel view and allows you to filter out exposures by dragging a selection box over an assortment of exposures points on any individual exposure plot (or by adjusting the slider or inputs above the plot). Clicking the “Select Thresholds” confirmation button will filter your exposures by rejecting any outside of these bounds. The threshold can be cleared by simply clicking the “Clear thresholds” button above the individual plot and controls.

<figure><img src="/files/S7ABIWnstJvdRqXtjr60" alt=""><figcaption></figcaption></figure>

Threshold selections are additive, and any exposures that do not fit within the combined selection will be rejected. Rejected exposures will show up as red points on all scatterplots (including the overview panel scatterplot) and will be differentiated in the table with red “Index” cells.

The other tabs in the action panel are “Individual” and “Micrograph”. In order to view data in these tabs you must first select an exposure, either through a scatterplot or the micrographs table.

* **Individual:** When an exposure is selected the “Individual” tab will be populated with a variety of information about that exposure, including its 2D CTF, Global Motion Trajectory, Local Motion Trajectory, and 1D CTF. Each of these detail panels within this view can be expanded by clicking the image, and the image can be downloaded by clicking the download button in the bottom right corner of each panel.
* **Micrograph:** The “Micrograph” tab houses an exposure viewer with functionality for zooming, panning, measuring, adjusting the lowpass filter, and downloading the image.

When an exposure has been selected, the rejection button in the top control bar can be clicked to individually reject that exposure.

Once thresholding and curation is complete, the “Done” button in the job control bar can be clicked to complete the job and output the relevant selections of exposures.

### Auto-Thresholds

In some cases it may be preferable to curate exposures based on a set of pre-defined thresholds. If one or more parameters within the *Auto Thresholds* section is set (numerical value range in the form of `min,max`), upon launching the job will skip the interactive process, apply the threshold automatically and generate relevant outputs. If an auto-threshold parameter is set but that attribute does not exist in the exposure input, the threshold will be skipped.

<figure><img src="/files/Uuhu2YtBtLt1f4vvxlBS" alt="" width="375"><figcaption><p>An Exposure Curation job in building status with an auto-threshold on CTF Fit Resolution (Å) set. When launched with this configuration, the job will output only exposures within a CTF Fit range of 0-4Å.</p></figcaption></figure>

## **Interactive Job: Select 2D Classes**

<figure><img src="/files/R68T0u7j1Mlt2dOtNRWN" alt=""><figcaption></figcaption></figure>

This job allows you to interactively select 2D classes from the output of a 2D Classification job. If particles (from the same 2D Classification job) are included in the inputs, then particles will also be partitioned based on their assignments to the selected class averages.

The Select 2D Classes job most commonly follows any 2D Classification job in which there are one or more "junk" classes. Removing "junk" particles from the particle dataset is an important step in particle curation, and will help increase the quality of 3D reconstructions and refinements. Note that Select 2D Classes can also be used on the output for any job that creates templates (e.g. on the templates generated in a Create Templates job).

### How to use Select 2D Classes

At the top right of the panel the number of currently selected particles as well as the total number of particles in the dataset are shown.

Each desired class can be selected by clicking on it. To improve the quality of your particle dataset, avoid selecting classes that contain only a partial particle, two or more particles, or a non-particle junk image (e.g. ice crystals). You can use both the number of particles and the provided class resolution score to identify good classes of particles. There are several ways to sort and filter the classes in ascending or descending order, shown along the top of the panel:

`Filter` : Filter the classes by selected or unselected.

`# of particles`: Sort by the total number of particles in each class.

`Resolution`: Sort by the relative resolution of all particles in the class (Å).

`ECA`: Sort by the number of Effective Classes Assigned (ECA).

In addition, for dealing with large numbers of classes, you may want to select all classes that meet a given threshold in particle count, resolution, or ECA value. To do this, first choose a class that has the desired threshold value. Next, right click on the chosen class or click the triple dot overflow menu button to reveal a context menu with options to select all classes that meet or fail this threshold.

The "Select All", "Select None", and "Invert Selection" buttons along the top of the panel can be used to quickly select all classes, clear all selections, and invert the current selection, respectively. These can be useful for dealing with a large set of class averages.

Filtering the classes by "Selected" or "Unselected" can be helpful to group all of these classes together in order to inspect them collectively and allow sorting on only these subsets.


# Upload Local Files

Upload files from your local computer to CryoSPARC directly in the browser

{% hint style="info" %}
The ability to upload local files is a new feature available in CryoSPARC v4.5 and later
{% endhint %}

In the course of a typical cryo-EM project, it is often necessary to view, generate or modify files using an external program, for example, downloading a volume from CryoSPARC, opening it in ChimeraX and using that program to create a mask which will ultimately be used in a CryoSPARC Local Refinement. These downloaded and modified files typically reside on a user’s laptop or local computer through which they are accessing CryoSPARC, and are therefore outside of the filesystem that CryoSPARC itself has access to.

To simplify the process of incorporating these external files into the cryo-EM workflow, CryoSPARC v4.5+ allows users to upload files from their local computer to the CryoSPARC system directly in the browser. The uploaded files are added to the CryoSPARC project directory.

## Uploading local files to CryoSPARC

The easiest way to upload files to CryoSPARC is to drag the file (or multiple files) you want to upload from your filesystem to any CryoSPARC window:

<figure><img src="/files/ZYxyunk2nKNMWON9kLlT" alt=""><figcaption></figcaption></figure>

This will open the Upload Files dialog, with the dragged file ready to upload to the CryoSPARC instance. If the file was dragged to a CryoSPARC window that does not currently have a Project open, a Project must be selected from the dropdown before proceeding:

<figure><img src="/files/5ctMCsQtdX5QAjmU6beR" alt=""><figcaption></figcaption></figure>

The Upload Files dialog provides several functions:

<figure><img src="/files/AdcWIasKK8fVXHSaHQww" alt=""><figcaption></figcaption></figure>

* Files uploaded to CryoSPARC through the browser are added to a directory named `uploads` in the selected CryoSPARC project directory. The upload dialog lists all files already in the uploads directory in the first panel (left in the image above).
* Clicking a filename in the uploaded files list copies the full path to that file to the clipboard (i.e., it is ready to paste as if it had been manually copied). This function is convenient for download or navigation in the terminal or interfacing with other programs.
* The second panel (right hand side in the image above) lists files which will be uploaded along with their size. Hovering over a filename will display an option to remove it from the list (meaning it will not be uploaded). Alternately, all files can be cleared from the upload list by clicking `Clear files`.
* More files can be added to the upload by dragging and dropping additional files into the Upload Files dialog, or by clicking `Upload additional files` and selecting them in the local filesystem.

Once the user clicks the green `Upload` button, the upload process will start. All files are uploaded in parallel. Any number of files can be uploaded simultaneously. Individual files larger than 10 GB are not currently supported.

Once the uploads have completed, the Upload Files dialog can be closed (by clicking the green Done button) or more files can be uploaded by dragging and dropping into the dialog.

### Other ways of accessing the Upload Files dialog

The upload dialog can also be opened by clicking Upload Local Files in the green New Job dropdown menu, under the Upload heading. This method also provides convenient access to the uploaded files list without having to upload a new file.

<figure><img src="/files/8UKwThaRW0uSxd69kcUJ" alt=""><figcaption></figcaption></figure>

The upload dialog is also accessible from the spotlight search, which can be accessed by clicking the magnifying glass icon in the left-hand sidebar or by pressing Ctrl + K (Cmd + K on Mac):

<figure><img src="/files/lfbzXFUnk0SC68DqeK7A" alt=""><figcaption></figcaption></figure>

<details>

<summary>How does the file upload work?</summary>

Individual files are split into 5 MB “chunks”, which are individually uploaded. During the upload process, the `uploads` directory will fill with these chunk files. Once the upload process is done, the chunks are combined back into the original file and removed.

If the upload process is interrupted, the chunks generated before the failure will be orphaned in the filesystem. They will not interfere with any subsequent uploads, and will be automatically removed if the file is successfully transferred later. It is harmless to remove these orphaned chunks.

</details>

### What file formats can be uploaded to CryoSPARC

The following file formats can be uploaded to CryoSPARC through the browser. Unsupported file types can be uploaded by first compressing them (e.g., to a `.zip` file) and uploading the compressed file.

<table><thead><tr><th>Item</th><th>File type</th><th data-hidden></th></tr></thead><tbody><tr><td>3D Volumes</td><td><code>.mrc</code>, <code>.map</code>, <code>.ccp4</code></td><td></td></tr><tr><td>Cryo-EM movies and micrographs</td><td><code>.eer</code>, <code>.mrc</code>, <code>.tiff</code>, <code>.tif</code></td><td></td></tr><tr><td>Cryo-EM particle stacks</td><td><code>.mrc</code>, <code>.mrcs</code></td><td></td></tr><tr><td>CryoSPARC metadata files</td><td><code>.cs</code></td><td></td></tr><tr><td>BILD files</td><td><code>.bild</code></td><td></td></tr><tr><td>Segmentation files</td><td><code>.seg</code></td><td></td></tr><tr><td>JSON files</td><td><code>.json</code></td><td></td></tr><tr><td>PDFs</td><td><code>.pdf</code></td><td></td></tr><tr><td>Text files</td><td><code>.txt</code></td><td></td></tr><tr><td>Compressed files</td><td><code>.tar</code>, <code>.tar.bz2</code>, <code>.tar.gz</code>, <code>.zip</code></td><td></td></tr><tr><td>Images and videos</td><td>Any file with the <a href="https://developer.mozilla.org/en-US/docs/Web/HTTP/Basics_of_HTTP/MIME_types">MIME type</a> image or video, including but not limited to <code>.jpeg</code>, <code>.png</code>, <code>.mp4</code></td><td></td></tr></tbody></table>

## Auto-Import 3D Volumes

In version 5.0 and later, adding a `.mrc` or `.map` file to the upload dialog automatically reveals a new table section that enables the creation and queuing of an *Import 3D Volumes* job for each file upon completion of the upload. This workflow is designed to streamline volume ingestion by integrating import and queueing directly into the upload process.

<figure><img src="/files/OekRhfzZhzc1Qc7nsqhW" alt=""><figcaption></figcaption></figure>

The workspace to which these jobs are queued is configurable and defaults to the current workspace if the upload dialog is opened from within one. Similarly, the lane used for queueing can be adjusted, with the master node selected by default. Selecting an individual file within the table toggles whether it should be automatically imported, while the checkbox in the table header allows this behavior to be enabled or disabled for all files simultaneously. The volume type for each entry can be set manually; however, it is automatically inferred and populated based on the file name when the file is added.

After selecting **Upload & Import**, the upload proceeds as usual. Once the upload completes, an *Import 3D Volumes* job is created and queued for each volume file with auto-import enabled. The job’s volume data path and volume type parameters are populated automatically using the file name and the volume type specified in the dialog.

<figure><img src="/files/kh4QADthF8bpmRPbBmq1" alt=""><figcaption></figcaption></figure>

## Name conflicts

CryoSPARC does not allow previously-uploaded files to be overwritten. If a file slated for upload has the same name as another file in the `uploads` directory, it will be highlighted with an orange background. The file with which it conflicts will have orange text.

<figure><img src="/files/63IgGsO5CTwKpI2Hgla1" alt=""><figcaption></figcaption></figure>

If a file with a conflicting name is uploaded, the new version will have the date (in YYMMDD format) and time (HHMMSS format) appended before the extension. The renamed file is highlighted in blue in the Upload Files dialog. The existing, previously-uploaded file remains unchanged.

<figure><img src="/files/oyPA4fF6heJZOhO7lyLE" alt=""><figcaption></figcaption></figure>


# Managing Data

{% hint style="danger" %}
Do not remove from the filesystem any directory that is managed by an[ attached CryoSPARC project](https://guide.cryosparc.com/application-guide/pages/F3KBgDxkuaoVRFwpV0KW#2.-attaching-detaching-archiving-and-unarchiving-projects). First delete unwanted projects using the *Delete Project* GUI action or the [`delete_project()` method of the CryoSPARC CLI](/setup-configuration-and-management/management-and-monitoring-4.7/cli-4.7#delete_project-project_uid-str-request_user_id-str-all_jobs_in_project-list-all_workspaces_in_projec).
{% endhint %}

Available in the management dialog, the Project Data and Session Data tabs allow you to view and perform actions related to managing data size.

## Project Data

The project data table is a comprehensive view of all available projects across the instance with the intention of allowing decisions to be made in regards to data size on disk. The table can be sorted by any of its fields and each project row houses a nested workspaces table. This nested table shows all workspaces within said project and pertinent information about them. The actions column makes available a “Refresh Project Stats” button for each row, intended to allow the fetching of the most up to date project sizes.

<figure><img src="/files/HIV62zhEhhPbFfbE5dJP" alt=""><figcaption></figcaption></figure>

## Session Data

The session data table is similar to the project data table in intention, with more scope in actionable options. The table shows all projects with available sessions as top level rows, and the contained sessions as collapsable sub-rows beneath. Sorting the table by any of its columns will sort both the projects and contained sessions by the attribute.

<figure><img src="/files/yzL0QiOIby3aaIF6QnIN" alt=""><figcaption></figcaption></figure>

The actions column allows the user to refresh stale data on the project level, which refreshes the project total size and the size of all sessions, or by a single session, which refreshes solely that session and the project total updated with its new size. The download button will download a set of sessions stats, and the link button will take you to the single session live view, closing the dialog.

Each cell in the session row for the columns Raw Data, Micrographs, Thumbnails, Particles, and Metadata, are interactive and will trigger a management menu when clicked. These menus give a variety of options for managing session data depending on the session status and row type.

<figure><img src="/files/WqGIGBNaxpFGWjuOphcV" alt=""><figcaption></figcaption></figure>

For more information about CryoSPARC Live Session Data Management, please see:

{% content-ref url="/pages/-MNexr5lrR2HBCuP4qWS" %}
[Guide: CryoSPARC Live Session Data Management (≤v4.7)](/setup-configuration-and-management/software-system-guides/cryosparc-live-session-data-management-4.7)
{% endcontent-ref %}


# Downloading and Exporting Data

## Downloading Lists of Projects, Workspaces, Sessions, and Jobs as a CSV File

The browse system includes the capability to download a CSV of the data shown in any particular view. This download will respect all filters and any sorting options applied to the view.

Initiating a CSV download can be accomplished by clicking the download button on the far right side of the application footer.

<figure><img src="/files/I9ZcBOStdCik2Te20IG9" alt=""><figcaption><p>CSV download button is located in the footer of all browse pages</p></figcaption></figure>

Pressing this button will open a dialog for customizing the CSV to suit your needs. Here you may update the CSV file name and select what information you would like to have appear in the CSV. The options shown in the “Table Columns” section represent columns in the CSV download and can be toggled for inclusion or exclusion. These options can also be dragged and dropped within the list to reorder the columns of the CSV table. Column options are mapped so that top to bottom in the download list corresponds to left to right in the CSV.

<figure><img src="/files/dUz9J88xXN7a2YQyt4dx" alt=""><figcaption><p>Dialog displaying options to customize the CSV file</p></figcaption></figure>

Clicking the blue “Download” button at the bottom of the dialog will download the formatted CSV to your device. In v5.0+ additional options have been introduced such as the inclusion of info tags and the ability to only download jobs that you have manually selected.

{% hint style="info" %}
You can download a list of all completed jobs across the entire instance this year by searching the spotlight (`command` + `k`) for ‘All jobs’ and adding a status filter of ‘completed’ and selecting a start and end date.
{% endhint %}

## Downloading Job Results

From the browse view, you can select a job and inspect it to view all job output groups. Each contain a list of items (such as `.cs` file metadata or `.map` volumes) to download:

<figure><img src="/files/XvhboRbxSdg8jFLwGQiZ" alt=""><figcaption><p>Viewing and downloading output results from the job inspection dialog.</p></figcaption></figure>

Alternatively, the list of output groups is available in the details sidebar of a job when selected under the “Outputs” panel:

<figure><img src="/files/VptGAGbjlg0Jlzhx4ZrE" alt=""><figcaption><p>Viewing and downloading outputs from the job details sidebar.</p></figcaption></figure>

Additional download options are available from the “Outputs” tab of the job inspection dialog. Here you can copy the file path or download individual low-level results as well as download results from a specific iteration:

<figure><img src="/files/Izmx91Oav2rp9gzw6I2t" alt=""><figcaption><p>Downloading the alignments3D low-level result from iteration 2 of this Homogeneous Refinement job.</p></figcaption></figure>

## Downloading the Job Event Log

Often it is helpful to download a standalone copy of the processing history of a job. In CryoSPARC v4.0 you can now generate a PDF job that contains a cover page of important metadata and the full event log including images. This makes it easy to archive and share the results of a job.

<figure><img src="/files/ZUP129WX3jI6dschU37r" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/XI0SxPzlJzXtGcJSaZ5U" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/Zp2vAxpJ6ICIzMEvRyAG" alt=""><figcaption></figcaption></figure>

## Downloading a Job Report

In addition to downloading just the job event log, you can download the event log and a set of CryoSPARC system logs for the purposes of debugging. You can choose to include or exclude images in the event log PDF. The report is packaged in a compressed ZIP file.

<figure><img src="/files/vqo9zSvSRXkpXum33Txl" alt=""><figcaption></figcaption></figure>

## Exporting Jobs

Refer to this guide on exporting jobs:

{% embed url="<https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/tutorial-data-management-in-cryosparc#use-case-share-a-particular-job-with-another-user>" %}

## Exporting Projects

Refer to this guide on exporting projects:

{% embed url="<https://guide.cryosparc.com/setup-configuration-and-management/software-system-guides/guide-data-management-in-cryosparc-v4.0+>" %}


# Instance Management

The management dialog can be opened on any page by clicking the “adjustments” button on the navigation bar located on the lefthand side of the app window. Once open, different sections can be navigated to using the tabs located at the top of the dialog, which correspond to sections for the management of jobs, tags, project and session data, as well as instance-level information. Alternatively, manage dialog sections can be navigated to directly by opening the navigation bar “triple dot” overflow menu and clicking on an option in the first section of the menu.

{% hint style="info" %}
The manage dialog will appear above the current page content without navigating you away from the page.
{% endhint %}

<figure><img src="/files/QaFI5yL05ahetMzT936X" alt=""><figcaption></figcaption></figure>

### Instance Tab

The instance tab provides a read-only view of all lanes and targets configured via the command line. For more information on how to configure CryoSPARC processing nodes, refer to the following page on the guide: [Connecting a Worker Node](https://guide.cryosparc.com/setup-configuration-and-management/how-to-download-install-and-configure/downloading-and-installing-cryosparc#connecting-a-worker-node).

<figure><img src="/files/suoqSmd9S8bHXzZMvigb" alt=""><figcaption><p>Instance Tab in v4.0</p></figcaption></figure>

v5.0 and later introduce a substantially redesigned *Instance* tab that expands both visibility and operational control. The updated interface provides a structured table of GPU information for available targets, along with comprehensive search and filtering capabilities across lanes, targets, and GPUs. It also enables direct copying of filesystem paths and cluster submission scripts, streamlining workflow integration. Lanes and targets can be viewed as cards or in a table.

<figure><img src="/files/IhVtzX91lmXSoTNr9Zn7" alt=""><figcaption><p>Instance Tab (Card View) in v5.0+</p></figcaption></figure>

<figure><img src="/files/AD1rANkO2sbrUEBCEhs9" alt=""><figcaption><p>Instance Tab (Table View) in v5.0+</p></figcaption></figure>

### Backups Tab

The backups tab displays a history of the most recent backups that were configured via the command line.

<figure><img src="/files/AZEmIJQZKOQLAkzWhi16" alt=""><figcaption></figcaption></figure>

### Notifications Tab

Notifications are presented in the CryoSPARC interface from local actions (such as queuing a job) and external events (such as the progress of a project export). In most cases these notifications are hidden when an event is resolved (i.e. the project export completes). In rare cases you may want to manually clear an active notification from displaying. This can be done by clicking on the ‘Clear’ button next to any active notification:

<figure><img src="/files/Eqzb06LJ0wCIOaw3VnQJ" alt=""><figcaption></figcaption></figure>

To view a history of inactive notifications, click on the ‘Inactive’ toggle:

<figure><img src="/files/bDgm2zEr80FvDBlkKChQ" alt=""><figcaption></figcaption></figure>


# Admin Panel

The admin panel can be accessed by clicking on the “key” icon located on the navigation bar on the lefthand side of the app window.

<figure><img src="/files/9dE3WHQxbO6jswHvnQhP" alt=""><figcaption></figcaption></figure>

## Instance Settings

<figure><img src="/files/VoMwjzJLm7ecH5s75EH6" alt=""><figcaption></figcaption></figure>

The instance settings tab allows you to set general instance wide defaults. Two configurable fields exist here, the instance name, and the default instance job priority.

* **Instance Name:** The instance name is simply a name to differentiate your instances visually.
* **Default Instance Job Priority:** This is the default priority that jobs run within the instance will be given when launched. The job priority differentiates the importance between different jobs launched in an instance and allows more important jobs to run before less important jobs. The priority can be set on a job by job basis, and on a user default basis. By setting the default instance job priority, users can be given priority above or below this mark, and jobs can similarly be launched above or below this mark.

If a “Lane Queue Override” has been set, an indicator will also appear here. This override allows for import jobs to be queued on lanes outside of the default.

## Instance Logs

<figure><img src="/files/0NrtX0Y3zDS0r5xkAbiH" alt=""><figcaption></figcaption></figure>

These are collections of the logged messages and errors from the instance. Selection menus on the top bar allow you to choose between a variety of different logs including: database, app, app\_api, app\_legacy, command\_core, command\_vis, and command\_rtp. The number of lines can be toggled to 50, 100, 200, or 500.

The download report button will download a zipped folder of all of the available log files as well as browser and runtime diagnostics files.

This information allows you to quickly diagnose any problems or behaviours your instance may be experiencing, and can be invaluable for debugging purposes.

## User Management

<figure><img src="/files/RbNRs2lvrzuEDdecgLF6" alt=""><figcaption></figcaption></figure>

The user management tab is the hub for curating users across an instance. User details are displayed and a variety of general actions are contained in applicable columns.

* **Email:** The user email can be edited by hovering the email cell and clicking the pencil icon. This will transition the email to an editable input. To exit without making any changes, simply click the “X” button, to confirm and update changes click the “check” button. (Clicking outside of the input will also close it without making any changes).
* **Role:** The user’s role can be changed between “User” and “Admin”. The admin privilege will allow that user to perform any admin actions across the instance, and will also allow them to see and manage other user’s work. This operation cannot be performed on the person currently editing, and so the selection button is disabled for the current user.
* **Tokens:** These are the reset and register tokens made available to allow a user to either reset their password, or register their account for the first time. An admin must deliver these tokens to the user so that they can use them on the register and reset pages.
* **Live Data Management:** This toggles the users ability to perform actions on specific sessions within the “Session Management Table” in the “Manage” dialog. If the user has this permission, they can perform action on the sessions such as deleting and archiving session data.
* **Job Priority Management:** This toggles the users ability to set a priority on a per job basis. If toggled on, the user will see the priority value field in the job builder, and can set it higher or lower than the instance or user priority value.
* **Default Job Priority:** This value represents the default job priority on a user to user basis. It will show the same score as the instance default if not modified, and can be set above or below the instance default to override that score on a per user basis.
* **Delete User:** Using this button will delete the user and all of their information from the instance. This is a non recoverable action, the deletion is immediate and permanent.

### Create New User

This form allows you to create a new user within the instance. The user will be immediately added to the instance and appear in the User Management table. They will need to be given the registration token code available in the table in order to set a password and gain access to their account.

<figure><img src="/files/BUZDLwzikTk1D5Hb9DJL" alt=""><figcaption></figcaption></figure>

### Cluster Configuration

The Cluster Configuration page allows you to set custom variables for cluster job submission scripts at an instance-wide and per-target level.

<figure><img src="/files/p5tfMPuuoPMjxtcrXInR" alt=""><figcaption></figcaption></figure>

For more information, see:

{% content-ref url="/pages/5zikU1RIs7uHyIpgV96G" %}
[Guide: Configuring Custom Variables for Cluster Job Submission Scripts](/setup-configuration-and-management/software-system-guides/guide-configuring-custom-variables-for-cluster-job-submission-scripts)
{% endcontent-ref %}

### User Lane Restrictions

Every user in CryoSPARC can be restricted to use only a subset of the lanes available in the instance. This page allows you to modify a user’s lane assignments.

<figure><img src="/files/DEZW1nzJdz8HXmx6dHKs" alt=""><figcaption></figcaption></figure>

For more information, see:

{% content-ref url="/pages/CJFjwLOpeqI6r91Mxkob" %}
[Guide: Lane Assignments and Restrictions](/setup-configuration-and-management/software-system-guides/guide-lane-assignments-and-restrictions)
{% endcontent-ref %}


# Keyboard Shortcuts

Keyboard shortcuts allow you to perform a variety of common actions more efficiently across the application.

The [Shortcuts Dialog](/application-guide/keyboard-shortcuts#shortcuts-dialog) shows all available keyboard shortcuts for the application inside of hierarchical sections in a single panel. It also includes a full “fuzzy” search to quickly locate a shortcut of interest.

<figure><img src="/files/iZEMLWgrbk0OC48u6Wtg" alt=""><figcaption></figcaption></figure>

The shortcuts dialog can be opened by clicking the “keyboard” icon button located on the main navigation bar at the leftmost side of the window. The button is located at the very bottom of the bar just above the user initials.

<figure><img src="/files/ADXFpeRUZN5wHJmVhsmu" alt=""><figcaption></figcaption></figure>

All of the keyboard shortcuts available across the application are included in this dialog. Shortcuts are located in sections based on the areas of the application where they are applicable. Sections are organized further into subsections to create clearer delineations of where the shortcut can be used. Each section can be collapsed to hide its contents, this will also collapse all of its subsections.

The **search bar** at the top of the dialog is designed to filter by section title, shortcut description, and the shortcut itself. Typing a search term into the bar will surface whichever one of these items that has the highest priority and closest match. If no matches are found, the search will attempt to surface shortcuts that are the closest match, or have related metadata. This can help you “fuzzy” search for shortcut that you may not remember exactly.


# Image Formation

The pages in this section provide background on the physical processes that occur in the microscope to generate and detect an image.

{% content-ref url="/pages/VIbGbRFssZDHvQpr4JM2" %}
[Contrast in Cryo-EM](/cryo-em-foundations/image-formation/contrast-in-cryo-em)
{% endcontent-ref %}

{% content-ref url="/pages/wMgsBSe2ZzNA0ekFngXN" %}
[Waves as Vectors](/cryo-em-foundations/image-formation/waves-as-vectors)
{% endcontent-ref %}

{% content-ref url="/pages/32GhFIW0DuCbnmEhQpxt" %}
[Aliasing](/cryo-em-foundations/image-formation/aliasing)
{% endcontent-ref %}


# Contrast in Cryo-EM

Where does contrast in an electron micrograph come from? What can we do to produce more contrast? What effects do we have to take into account when processing data from electron micrographs?

## Contrast

Contrast is a general term for differences in a signal (typically, an image). For the purpose of this guide, it is sufficient to think of contrast as follows:

{% hint style="info" %}
**Contrast** is the difference in intensity (i.e., brightness/darkness) between the darkest object and lightest object in an image
{% endhint %}

For example, consider this image:

<figure><img src="/files/jlSeVuSAHGXg6Zov0iSu" alt="In image, a black square and a white square stand on a grey field. A colorbar at the right tells us that the pure white square has a value of -1, while the pure-black square has a value of +1." width="506"><figcaption></figcaption></figure>

The color bar on the right shows the true value of each pixel — you could think of this value as the electron density, for example. The image above has very good contrast: the darkest region is pure black, and the lightest region is pure white.

This same image with poor contrast might look like this:

<figure><img src="/files/H54TwdC0ELPJxmaBiWS2" alt="This image shows the same squares as the previous image (one with a value of +1 and one with a value of -1). However, these squares are not pure black or white anymore, but a darker or lighter shade of grey." width="506"><figcaption></figcaption></figure>

Note that the *actual values* of each square and the background are the same. However, the intensities in the image are closer to each other, which makes the *differences* in intensity smaller. Contrast is a property of the *image*, not of the *objects*. Of course, the process by which we record the image and properties of the objects do have an impact on the image’s contrast. The rest of this page investigates the process of contrast generation in the specific context of cryo-EM.

### Image formation in Cryo-EM

Cryo-EM micrographs are formed when an electron beam passes through the sample and hits the detector. The detector counts the number of electrons that hit each pixel over the duration of a frame and stores that value as the brightness of that pixel.

If there were no sample, the beam would travel through the microscope and hit the detector. In a well-aligned microscope, this means that each pixel would receive the same number of electrons. There would therefore be a flat grey micrograph, with each pixel only differing by a small amount due to noise and random chance:

<figure><img src="/files/9qoPeh5P8wNPNmGUzGAj" alt="A very simplified cartoon animation of a theoretical electron beam. Rays travel from a silver object at the left of the frame and hit an array of pixels at the right. As beams hit pixels, they become brighter. At the end of the animation, almost all of the pixels are lit. Some are brighter than others."><figcaption></figcaption></figure>

If there is a sample between the beam and the detector, the sample will block the electrons. If we assume that no electrons which pass through a sample, we would have near-perfect contrast:

<figure><img src="/files/51wCS0UIxskM6RMFzgRG" alt="Another cartoon animation of an electron beam. In this one, a cube and sphere stand between the detector array and the electron beams. The detector pixels behind these objects are never lit up, because the beams are blocked by the objects. The other pixels are mostly lit up, with the notable exception of a pixel near the sphere&#x27;s shadow, which by chance is never hit by an electron beam."><figcaption></figcaption></figure>

In reality, contrast in cryo-EM is much more complex than this. Electrons almost always pass directly through biological samples without being blocked or absorbed. Furthermore, both the wave and particle nature of electrons in the microscope must be considered in order to understand contrast completely. This guide page covers the mechanisms by which electrons interact with biological samples, and the types of contrast these interactions introduce into the images.

### Amplitude contrast

{% hint style="info" %}
**Summary:** Amplitude contrast occurs when an object directly reduces the magnitude of the incoming wave.
{% endhint %}

Perhaps the most obvious type of contrast is amplitude contrast. When a wave (e.g., the electron wave) passes through an object, some amount of that wave will be absorbed, deflected, or otherwise blocked by the object. Thus, the exiting wave has a lower amplitude. Since image intensity is the square of the amplitude of the wave exiting the sample, a change in amplitude results in contrast between the object and its surroundings.

<figure><img src="/files/hpaubazadkWWtJQjxfJS" alt="Two yellow sine waves are depicted in this animation. The bottom sine wave travels left to right without encountering any sample. Its amplitude stays the same the whole way across. The top sine wave encounters a glass cube. As the wave travels through the cube, its amplitude smoothly decreases until it leaves the cube with an amplitude approximately 1/5 that of the bottom wave."><figcaption></figcaption></figure>

In the animation above, we compare a wave which passes through a sample (top) to a wave which directly hits the detector (bottom). Each wave’s amplitude is marked with a yellow bar on the far right side. This sample has significant amplitude contrast — the wave’s amplitude is reduced by 80% after it exits the sample.

The heavy atoms used for negative stain electron microscopy deflect a large proportion of the incoming electron beam, giving them excellent amplitude contrast. However, amplitude contrast plays a negligible role in cryo-EM of biological samples.

Most biological macromolecules are composed of atoms with similar atomic number to those of the aqueous buffer they’re suspended in. This means that both the “empty” ice and the protein both block electrons with roughly the same strength, so there is little amplitude contrast between the two. A notable exception is nucleic acids, which have heavier phosphorus atoms and so they have a slightly higher amount of amplitude contrast.

The vast majority of contrast in the case of cryo-EM comes, instead, from phase contrast.

### Phase contrast

{% hint style="info" %}
**Summary:** Phase contrast occurs when an object delays or advances the phase of the incoming wave. It is not directly detectable since the amplitude of the wave is unchanged.
{% endhint %}

Proteins belong to a class of objects called “phase objects”. Instead of absorbing or blocking the incoming wave, phase objects *delay* the wave, resulting in a phase shift between the incoming wave and the exiting wave.

Although the phrase “phase object” may be unfamiliar, many objects in science or even day-to-day life are phase objects. For instance, glass diffracts light by delaying its phase, and cells have poor contrast in bright-field light microscopy because they are phase objects.

<figure><img src="/files/zm4qq16IEvblTCX3wb2M" alt="Two blue sine waves travel from left to right in this animation. The peaks of each wave are connected with a straight line. The bottom wave encounters no object, and so travels without changing. The top wave travels through a glass cube. The peaks become squished together as the cube shifts the wave&#x27;s phase. The lines connecting the peaks become diagonal as the peaks in the top wave are slowed by their phase shift."><figcaption></figcaption></figure>

In the above animation, the glass cube is now a pure phase object. Phase objects delay or advance the phase of waves which pass through them. They do not, however, change the amplitude of the incoming wave. Note that at the left hand side, the waves are in sync. However, the top wave is delayed while it travels through the cube. The end result is that when the waves encounter the detector at the right hand side, their amplitude is unchanged, but their phase is shifted by 180°, meaning the peaks line up exactly with the troughs.

Detectors, however, cannot detect phase. They record intensity, which is the square of the wave’s amplitude. When these waves hit the detector, their amplitudes are identical. Thus, normal images of pure phase objects have zero contrast — they are completely invisible!

Given the above, it is natural to ask: how can we produce detectable contrast for a phase object in the microscope? This kind of contrast is called phase contrast, and can be produced by setting up the microscope in a way that causes the incoming wave and delayed wave to interact via interference.

In the above animation, note that the peaks of the top wave align perfectly with the troughs of the bottom wave. When these waves are added together, this will produce a net amplitude of 0. In this way, the phase object has created detectable amplitude contrast in the final image!

Phase contrast is explored in much greater detail in later sections of this page.

### Weak phase object approximation

{% hint style="info" %}
**Summary:** Weak phase objects scatter only a small proportion of the incoming wave. This lets us approximate the (otherwise highly complicated) exit wave as the incoming wave plus a small, scattered wave with a π/2 phase shift. **Critically, under this approximation, the amplitude of the scattered wave is linearly proportional to the object's density**. In other words, the scattered wave carries an image of the particle.
{% endhint %}

Real cryo-EM samples produce very small phase shifts in the electron beam, typically on the order of tenths of a degree.

{% hint style="info" %}
If the vector representation of waves is unfamiliar, it is explained in the [Waves as Vectors](/cryo-em-foundations/image-formation/waves-as-vectors) article.
{% endhint %}

When viewing a single wave like this, it is easy to imagine directly modeling the phase shifts caused by the sample. However, recall that in the real world a 3D sample is imparting an intricate phase shift on a plane wave:

<figure><img src="/files/udPCusfGjzSAzdt2pMCK" alt="At the start of this animation, a cartoon of a protein is visible in the right-hand size. A flat white square is visible in the top left and a glowing square is visible in the center. The glowing square (representing a plane wave) travels to the right. This glowing square represents the point at which a plane wave has amplitude 1. As it passes through the protein, parts of the wave are stretched behind it, since those parts are phase shifted. The white square shows the same displacement as the glowing square, allowing us to inspect the deformation more closely."><figcaption></figcaption></figure>

In the animation above, the plane wave travels from left to right. As it encounters the sample, parts of the wave are delayed as they travel through regions with differing potential (a stationary copy of the plane wave is visible in the top-left to make the phase shifts clear). Remember that these phase shifts are invisible to the detector.

The exact nature of these phase shifts is extremely complicated, depending on a great number of properties of the scattering object and the incoming wave. Perfectly modeling even a single image formed by this process is computationally impractical, and intractable for the hundreds of thousands or even millions of images that need to be modeled during a single 3D map refinement.

Instead of calculating the exit wave explicitly, it is common in cryo-EM to use the *weak phase object approximation*. In the *weak phase object approximation*, we assume that the sample only scatters a small proportion of the incoming wave, and that the scattered wave has a constant phase shift of exactly π/2. The exit wave is therefore modeled as the incoming wave plus this small, π/2 shifted wave.

{% hint style="info" %}
Many derivations of the Weak Phase Object Approximation are available in the literature, typically incorporated into derivations of the CTF. We have included links to several in the References section at the end of this article.
{% endhint %}

<figure><img src="/files/YiakhLySwGD4wdWZaOhT" alt="This image has two parts: Real Wave and Weak Phase Object Approximation. In the Real Wave part, two vectors are labeled Scattered wave and Incoming wave. The two vectors are related by a phase shift of theta. In the Weak Phase Object Approximation part, the same two vectors are labeled Exit wave and Incoming wave. A third wave is labeled Scattered wave and points from the tip of the Incoming wave to the tip of the Exit wave. It has a phase shift of exactly one half pi and a length of theta." width="563"><figcaption></figcaption></figure>

Again, the utility of the Weak Phase Object Approximation is difficult to grasp when considering a single vector. However, again considering a plane wave traveling through a 3D object, its usefulness is apparent:

<figure><img src="/files/Pbi4Ryqx9SZ66mV8hJJt" alt="In this animation, the same protein and glowing plane wave are displayed as before. However, on the left side of the view there are now two squares. At the top is a black square representing the scattered wave. At the bottom is a white square representing the incoming wave. As the plane wave travels through the sample, an image of the protein appears in the scattered wave (linearly related to the sample density at each point). The incoming wave is unchanged."><figcaption></figcaption></figure>

In this animation, the plane wave again travels left to right. Under the weak phase object approximation, the phase shifts in the incoming wave are modeled as a scattered wave (top) and an unscattered wave (bottom). The scattered wave has a uniform phase shift of π/2, but more importantly, it has an amplitude that is *linearly proportional to the scattering potential of the sample*. If we could recover an image of the scattered wave alone, we would have an image of the sample.

However, the wave that reaches the detector (the exit wave) is the sum of the scattered and unscattered wave. The scattered wave’s phase shift of π/2 and relatively minuscule amplitude means that the magnitude of the exit wave is unchanged by the scattered wave. Thus, even under the weak phase object approximation, phase objects are still invisible in images!

With the weak phase approximation in place, we can finally discuss the mechanisms that cause signal from the scattered wave to appear in the final image as contrast. There are two primary mechanisms:

* Making use of imaging aberrations that produce additional phase shifts of the scattered wave. This is done by intentionally collecting data out of focus.
* Introducing additional phase shifts to the scattered wave but *not* the unscattered wave. This is achieved using a phase plate, which may be simplest to understand for readers who are familiar with optical microscopes.

## Electron-optical aberrations

### Defocus

{% hint style="info" %}
**Summary:** Focusing the microscope on a plane some distance away from the sample introduces contrast because scattered waves must travel further than the unscattered wave, causing a phase shift proportional to the scattering angle and distance from the sample.

When the scattered wave is added to the unscattered wave, this phase shift causes a difference in amplitude, which can be detected as an image.
{% endhint %}

So far in our discussion, the weak phase object approximation indicates that, under standard imaging conditions, images of proteins at focus would have no contrast. What happens if we collect an image at some distance away from focus? To answer this question, we need to consider the fact that the scattered beam changes direction after interacting with the sample.

The scattered wave is actually scattered in several directions. Each spatial frequency in the object scatters the electron beam at the angle $$\theta{} = \arcsin{\frac{\lambda{}}{2d}}$$, where $$\lambda$$ is the electron beam’s wavelength and $$d$$ is the spacing (the spatial resolution) of the object feature in question. For instance, all of the 10 Å information in an object scatters the electron beam at an angle of

$$
\theta = \arcsin{\frac{1.97 \times10^{-12}\ \mathrm{m}}{2 \times (10 \times10^{-10}\ \mathrm{m})}} = 9.85\times10^{-4}\ \mathrm{radians}
$$

Note that this scattering occurs in both the +θ and -θ directions, but we only show one of the two beams here for readability. This relationship is called Bragg’s Law and is important in many imaging fields, including x-ray crystallography.

<figure><img src="/files/PIthM2h5A0oy7g8bg4HB" alt="An animation of scattering.  Ten points are visible in the center of the image. They are evenly spaced in a row. The space between two points is labeled &#x22;d&#x22;.  Horizontal glowing lines, representing plane waves, enter from the top of the screen. When they hit the points, the lines continue downward but a new set of lines appears moving at an angle theta relative to the initial waves.  The spacing d between points shrinks during the animation. As it does, the scattering angle widens."><figcaption></figcaption></figure>

Since the beams are traveling in slightly different directions, they take slightly different amounts of time to reach the same position, which means the phase of the scattered beam shifts relative to the unscattered beam.

<figure><img src="/files/ruBXBUVwQpYHodQHORpd" alt="At the left, two lines (representing the scattered and unscattered waves) are shown at some angle theta. A pink plane indicates the point at which we&#x27;re sampling the two waves (the very top in this image). In the center, we see the unscattered wave and scattered waves as well as their sum. Since there is no defocus here, the scattered wave has a phase shift of half pi and does not significantly change the exit wave. The waves are represented as vectors at the right."><figcaption></figcaption></figure>

In the image above, the pink plane on the left represents the plane of the image. At the sample, the unscattered (yellow) and scattered (blue) waves have traveled the same distance (that is, no distance). Their path length difference is therefore zero, so the difference in their phases is unchanged. The scattered wave has a phase shift of π/2 relative to the scattered wave. Thus, the sum of the two waves (pink) looks identical to the unscattered beam. The sample is invisible at focus.

As the waves travel through space, the scattered wave must travel an additional distance the same vertical position as the unscattered wave — their path lengths are different to reach the same point in space. The further from the sample we measure their phases and the higher the spatial frequency which scattered the beam, the more the path lengths differ.

<figure><img src="/files/9ejlnAHxq7IdPuM8CILt" alt="The same image as above is animated. As the pink plane moves from the top to the bottom, the path length difference increases. This is visible as the scattered wave&#x27;s phase shifting, and therefore the exit wave&#x27;s amplitude as well."><figcaption></figcaption></figure>

As we collect an image further and further from focus, the path difference (and therefore phase difference) between the scattered and unscattered beams increases. When the path length differs by a quarter of a wavelength, the scattered beam has a phase shift of π. This means it is destructively interfering with the scattered wave, and the image amplitude is decreased (negative contrast). When the path length differs by three quarters of a wavelength, the opposite happens — the phase of the scattered and unscattered beams are the same, so the image amplitude is increased (positive contrast).

<figure><img src="/files/6DfRtFqSJshVhp3cdPMi" alt="Scattered, unscattered, and exit waves, represented as vectors, are shown without an additional phase shift and with an additional phase shift of half pi or three-halves pi. Without an additional phase shift, the scattered wave has a total phase shift of half pi and so does not produce appreciable amplitude contrast. With half pi additional phase shift the scattered wave is pointing in exactly the opposite direction of the unscattered wave, and so produces negative contrast. The opposite happens with three-halves pi additional phase shift, producing positive contrast." width="563"><figcaption></figcaption></figure>

Note that this phase shift depends on both the distance from the sample *and* θ, which itself depends on the spatial frequency d. We can thus expect the phase of the waves carrying information from the sample to differ based on both

* the distance away from the sample at which we image those waves (that is, the defocus), and
* the resolution of the information that initially scattered those waves.

To observe this effect, we can plot the contrast at several fixed distances from the sample with varying spatial frequency. In the plot below we plot the contrast between the scattered and unscattered waves at varying distances from focus (indicated on the right in microns). The X axis represents increasing spatial frequency. Contrast is plotted on the Y axis. When a line is at zero, the scattered wave has a phase shift of π/2 or 3π/2, so it does not interfere with the unscattered wave and there is no contrast. When a line is above zero, the scattered and unscattered waves are interfering constructively, generating positive contrast. When a line is below zero, the scattered wave is interfering destructively, generating negative contrast.

<figure><img src="/files/v5fMhNgYcYDCLtcCmYqL" alt="A plot of contrast at varying defocus and frequency values. The graph without any defocus is a constant line at 0 contrast. The others oscillate with increasing frequency between +1 and -1."><figcaption></figcaption></figure>

Each frequency of information present in the sample scatters a wave at a different angle. We visualize these waves in the X axis, with the lowest frequency waves at the left and the highest frequency waves at the right.

Because they have a different scattering angle, their path length also changes as they travel down the microscope. At 0 µm from the sample, all of the scattered waves are shifted by π/2 compared with the incoming wave, so there’s no contrast. The further they get from focus, the more their path length changes, and so the more their relative phase difference increases. Contrast increases the closer a wave’s phase shift is to an integer multiple of π; we see these maxima and minima in the Y axis. **Note that it is not possible to collect an image in which&#x20;*****all*****&#x20;of the waves have contrast!**

Note also that the very lowest frequencies essentially never change from 0 contrast. Their scattering angles are so small that their path length differences are negligible.

Path length difference is the fundamental source of contrast when imaging out of focus: instead of collecting an image focused on the specimen plane, where all spatial frequencies have a phase shift of π/2, we collect an image offset from that plane by some distance. This image has contrast in some spatial frequencies (and not others!) due to path length differences caused by the scattering angle and amount of defocus.

### Spherical Aberration

{% hint style="info" %}
**Summary:** Spherical Aberration describes the fact that electron microscope lenses overfocus rays that are further from the optical axis. This introduces an additional phase shift to waves scattered by high-frequency features, irrespective of defocus. Spherical Aberration therefore is another mechanism by which the scattered wave becomes visible as contrast in an image.
{% endhint %}

Although the curves describing the effect of defocus above may look familiar, they are missing a correction for an imperfection in the electromagnetic lens called spherical aberration. Lenses in the electron microscope focus waves more strongly the closer they are to the edge of the lens.

<figure><img src="/files/j6WYKzVuk8ojbLaZss13" alt="Several rays are emitted from two points on the left-hand side of the image. The rays from each point are focused by a lens. In the top copy, labeled &#x22;Ideal Lens&#x22;, the rays all focus on a single point. In the bottom copy, labeled &#x22;Spherical Aberration&#x22;, the rays focus on different points depending on their scattering angle."><figcaption></figcaption></figure>

Since higher-frequency features scatter with a greater angle, this adds a difference in focus to waves depending on their scattering angle, which also changes their path length. Spherical aberration generally starts to affect data only at higher resolutions, since it scales with the fourth power of the spatial resolution. We can add a term to reflect this to the function above to model spherical aberration in the sample.

<figure><img src="/files/rachA27TM8gVyE30YI3g" alt="The same defocus plot as above, but taking into account spherical aberration. Now even the 0 defocus graph has a slow oscillation between +1 and -1 contrast."><figcaption></figcaption></figure>

At low defocus, the most obvious effect of considering spherical aberration is the introduction of contrast without any defocus. Since spherical aberration scales with the fourth power of scattering angle, it does not begin to change the path length until higher resolutions, but then oscillates quite rapidly. This is more apparent when considering resolutions up to 1 Å:

<figure><img src="/files/OJl9gfuJdBhoB4iynOzL" alt="A third and final copy of the CTF at various defocus, but showing all the frequencies out to 1 angstrom. The 0 defocus plot does not start appreciably oscillating until around 3 angstroms and higher."><figcaption></figcaption></figure>

### Phase Plates

It is also possible to produce contrast by passing the scattered beams (but *not* the unscattered beam) through an environment that induces an additional phase shift (called a “phase plate”), ideally of π/2 or -π/2. The most commonly used phase plates are made of amorphous carbon, but others are under development. Phase plates are currently not often used in single particle analysis.

## The 1D CTF

With the weak phase approximation, defocus, and spherical aberration, we can write down the classic equation for the CTF:

$$
\mathrm{CTF}(f) = \sin(-\pi\lambda{}\delta{}f^2 + \frac{\pi}{2}C\_s\lambda{}^3f^4 + \phi{})
$$

where:

* is the spatial frequency (e.g., $$0.2\times10^{10}$$ for 5 Å information)
* $$\lambda{}$$ is the electron wavelength
* $$\delta{}$$ is the defocus (i.e., distance *above* the sample we focus the image on)
* $$C\_s$$ is the spherical aberration constant (a property of the microscope)
* $$\phi{}$$ is the additional phase shift introduced by a phase plate (if no phase plate is used, $$\phi{} = 0$$)

This equation allows us to compute the value of the CTF (a number between -1 and 1, telling us how much contrast is present) at each spatial frequency.

## The 2D CTF

The equation above only models phase shifts and scattering in a single direction. In reality, our images result from plane waves, meaning that the phase shifts and contrast are a function of both X and Y.

<figure><img src="/files/cuKU8nS2X6URaz63MWYa" alt="" width="563"><figcaption></figcaption></figure>

A one dimensional equation is no longer sufficient to represent these CTFs — we must consider both scattering direction and angle. Thus, the CTF plots become two-dimensional. You can imagine the 1D CTF graph rotating around the Y axis to generate a 2D surface. The highest points on the surface correspond to the positions with greatest positive contrast, while the lowest points correspond to the greatest negative contrast. These regions are often colored white and black, respectively, creating a 2D image.

<figure><img src="/files/CLp2hSkQrr7uz7RwAgG2" alt="At left, an oscillating CTF function is shown. When the contrast is +1, the line is white. When the value is -1, the line is black. This line is rotated 90 degrees about the horizontal axis. We now look at a rectangle with black where the CTF is -1 and white where the CTF is +1."><figcaption></figcaption></figure>

<figure><img src="/files/bEeInHBQE7K4hsmBpf0q" alt="The CTF slice from the previous image is rotated about the center, creating concentric rings of light and dark values. These pixels represent the contrast produced by the CTF for a given frequency and scattering direction." width="563"><figcaption></figcaption></figure>

In a 2D CTF, each pixel represents a particular wave. The pixel’s angular position denotes the wave’s orientation, while the pixel’s distance from the center of the CTF determines its frequency. Put another way, all pixels in a circle of a given radius represent waves of the same frequency, while all pixels on a line that passes through the center of the image represent waves diffracting in the same direction.

There is no guarantee that these plane waves are perfectly symmetrical. In fact, all cryo-EM images contain some amount of asymmetry due to anisotropic electromagnetic fields in the microscope lenses. This anisotropy is called astigmatism, and results in an under- or over-focused beam depending on scattering direction. This in turn results in an image with different contrast depending on the direction of the scattered wave.

<figure><img src="/files/UITh9sZBoejpsxTcHpW8" alt="Two CTF plots are shown. On the left (labeled &#x22;No astigmatism&#x22;), the oscillating contrast rings are circular. On the right (labeled &#x22;Astigmatism&#x22;), the oscillating rings are elliptical."><figcaption></figcaption></figure>

Minor astigmatism is observed in almost every dataset. It is thus important to fit the CTF in two dimensions rather than just the 1D rotational average. All CTF fitting algorithms in CryoSPARC fit astigmatism.

#### Higher-order aberrations

Higher order aberrations can also affect the CTF, but are typically only important at very high resolutions. These aberrations can be corrected using [Global CTF Refinement](/processing-data/all-job-types-in-cryosparc/ctf-refinement/job-global-ctf-refinement) and are discussed in more detail in the [CTF Refinement tutorial](/processing-data/tutorials-and-case-studies/tutorial-ctf-refinement).


# Waves as Vectors

{% hint style="info" %}
**Summary:** Waves can be represented by vectors. The length of the vector represents the wave's amplitude and the direction the vector points represents the wave's phase. The wave's oscillation can be represented by rotating the vector.
{% endhint %}

## Waves as vectors

A wave can be described by its amplitude, frequency, and phase. Phase describes how a wave evolves through time. As a quantity, it generally only makes sense as a phase shift relative to some other wave. For instance, these two waves are shifted by a quarter of their wavelength.

<figure><img src="/files/UXD1tUYPyY1Tw8EDKTf2" alt="A single cycle of two sine waves is shown. They are shfited by pi/2 radians, so the peak of one wave lines up with another wave crossing zero."><figcaption></figcaption></figure>

We would describe this as a phase shift of 90° or, more typically, π/2 in radians. At first, the notion of describing a phase shift (which in this graph looks like a movement of the wave left or right) as a rotation may be confusing. It can be helpful to imagine these waves as three-dimensional helices viewed from the side, rather than 2D waves:

<figure><img src="/files/qoUA9fRXfzvCf7pIlQyE" alt="At the top, a sine wave oscillates up and down. A white line is drawn across the zero point of the wave, and a white arrow points up or down to the value of the wave at the furthest-right point. Bottom left: the same wave is displayed as a helix, viewed from an oblique angle. The arrow points to the tip of the helix. Bottom-right: the helix is viewed directly down its helical axis. The wave now looks like a circle, with the arrow rotating smoothly in a circle."><figcaption></figcaption></figure>

Here, we see that a sine wave can be modeled as the rotation of a vector (blue arrow, right of the animation) with a length equal to the wave’s amplitude at a rotational velocity of $$\omega = 2\pi{}f$$, where $$f$$ is the wave’s frequency. Put another way, the speed of rotation represents frequency — a vector which spins faster traces out a wave which oscillates more frequently.

Now, consider what happens when we rotate the vector by π/2:

<figure><img src="/files/sQI5ifRgy78l6yxqHT7C" alt="This animation is similar to the previous one, except a pink wave has been added with a pi/2 phase shift. In the bottom-right, the arrows make a right angle as they rotate in a circle."><figcaption></figcaption></figure>

The pink wave appears to be shifted forward in time compared to the blue wave by a quarter of the wavelength.

### Adding waves

{% hint style="info" %}
**Summary:** Adding two waves together can change their amplitude, phase, or both.
{% endhint %}

The vector representation is especially useful when we begin to consider sums of waves rather than individual sine waves. For instance, it is an intuitive result that when we add two sine waves with the same frequency and phase together we get another wave of the same frequency and phase, but with a greater amplitude:

<figure><img src="/files/AUOfdTnq0ZXxSKCv6V0y" alt="Left: two sine waves of the same frequency and phase but differing amplitudes. Right: the waves added together have the same frequency but a greater amplitude."><figcaption></figcaption></figure>

Using the vector notation for this simple example, we observe the same behavior:

<figure><img src="/files/PVXGP4tlQqFEXzAO9hP5" alt="An animation of the addition above. A short blue arrow and a short red arrow are rotating around their bases. When the waves are added, they form a longer line which is also rotating around its base."><figcaption></figcaption></figure>

Adding the two vectors together produces a vector pointing in the same direction, rotating at the same speed, but with a longer total length. In this case, the utility of thinking about adding waves this way may be unclear, but consider the following surprising result:

<figure><img src="/files/N0RGN3ZUZZUJtZlQ5akf" alt="Adding a large wave and a small wave with a pi/2 phase shift results in a wave which looks like a phase-shifted copy of the first wave."><figcaption></figcaption></figure>

In this example, the red wave has the same frequency as the blue wave, but a much smaller amplitude and a π/2 phase shift. When we add together these two waves, we get the purple wave as a result — it looks like a phase-shifted version of the blue wave! This result is less surprising if we represent the waves as vectors instead:

<figure><img src="/files/re1NFOBZvTl9e7xD9NS5" alt="Adding a small wave with a pi/2 phase shift results in a vector which follows the hypotenuse of the triangle formed by the two waves. When the phase-shifted wave is small, this resulting wave has essentially the same magnitude as the original wave."><figcaption></figcaption></figure>

Because the red vector is always pointing perpendicular to the blue vector, adding the two has the effect of rotating the blue vector with only a very modest effect on the final magnitude. Recalling that a rotation is equivalent to a phase shift, we have arrived at the same result as directly adding each point of the wave.

The animation of vector rotation is helpful for developing a sense of what these vectors represent, but makes the figures cumbersome. For the rest of this guide, we will only draw the waves at some static position — implicitly, the vectors rotate as time passes or, equivalently, as we move through space. For instance, the above animation would be drawn like so:

<figure><img src="/files/kTjX4d931ulfE5kXzqsh" alt="A right angle formed by three vectors, labeled Wave 1, Wave 2, and Result (hypotenuse)."><figcaption></figcaption></figure>

Using this method, it is clear that the resulting wave has approximately the same amplitude as wave 1 (precisely $$\sqrt{|Wave\ 1|^2 + |Wave\ 2|^2}$$, where $$|Wave\ 1|$$ is the magnitude of Wave 1’s vector), but has a phase shift of $$\arctan{\frac{|Wave\ 2|}{|Wave\ 1|}}$$. If $$|Wave\ 2| \ll |Wave\ 1|$$, we can approximate the Result wave by shifting the phase of Wave 1 by the magnitude of Wave 2. This approximation is closely related to the Weak Phase Object approximation, covered in [Contrast in Cryo-EM](/cryo-em-foundations/image-formation/contrast-in-cryo-em).


# Aliasing

{% hint style="info" %}
**Summary**: When a signal is sampled at some sampling rate, the maximum frequency that sample can represent is twice the sampling rate. In images, the sampling rate is the pixel size. Thus, the pixel size sets an upper bound on the best resolution achievable for a given set of images.
{% endhint %}

## Aliasing

In the physical world, signals are continuous. The pressure waves which make up sound vary smoothly between high and low pressure; light waves which form images vary from light to dark, etc. However, computers must represent these signals with discrete samples. When we sample a signal, we take measurements of that signal at evenly spaced positions in space or time.

For a concrete example, consider the samples below:

<figure><img src="/files/OqSDqK6ibTeFoKKy76Vx" alt="Seven dots, clearly arranged along a single cycle of a sine wave"><figcaption></figcaption></figure>

By eye, it seems obvious that these samples come from a simple sine wave which oscillates once:

<figure><img src="/files/DhD3sMj2ianIoWBzEetO" alt="The seven points are connected by a sine wave which oscillates once"><figcaption></figcaption></figure>

However, these samples are explained just as well by a wave which oscillates seven times:

<figure><img src="/files/h1XR0CHAG5P4UapXoKov" alt="The same seven points are now connected with a sine wave which oscillates seven times"><figcaption></figcaption></figure>

or thirteen times:

<figure><img src="/files/Au5XhJ4jNHl7Es4jh98z" alt="The same seven points are connected by a wave which oscillates thirteen times"><figcaption></figcaption></figure>

Indeed, there are in fact an infinite family of waves which perfectly fit *any finite set of samples*. There is no way of knowing which of these infinite waves truly gave rise to the samples we observe, so by convention we select the lowest-frequency wave which fits the observed data. However, what happens when there really *is* high-frequency information in our images which our samples cannot capture?

### Nyquist Frequency

Say you collect an image of some object using a camera sensor which is 12 Å wide and has 6 pixels. Your pixel size is therefore 2 Å. Put another way, you are sampling the incoming image with a *sampling rate* of $$\frac{1}{2 \AA}$$ , or one sample for every two Å. If your object is a sine wave which oscillates with a frequency of $$\frac{1}{12 \AA{}}$$, the image is unambiguous:

<figure><img src="/files/GMZ6YUT0NslEqTbkUqRr" alt="The same set of points are shown twice. At top, they are connected by a sine wave which oscillates once. On the bottom, they are connected by steps at the height of each sample, and a dotted wave which oscillates once."><figcaption></figcaption></figure>

There are plenty of samples along the entire wave to accurate capture its shape. What happens when the wave’s frequency increases?

<figure><img src="/files/4B1y89biThGUoLVN2Kk5" alt="The wave at the top now oscillates twice, and the steps at the bottom also oscillate twice. The dotted wave is correct, also oscillating twice."><figcaption></figcaption></figure>

When you image an object with a wavelength of 6 Å, the result is still correct. The exact position of the object relative to the pixels has imposed some asymmetry in the image, but if the object shifted left or right the result would again be symmetric.

What about an object with even higher frequency?

<figure><img src="/files/O6HSFFVH0rr1MDwqwrNw" alt="The wave at the top now oscillates four times. At the bottom, the steps only oscillate twice, so the dotted wave only oscillates twice as well. The phase of the steps and dotted wave are also flipped."><figcaption></figcaption></figure>

When the object has a wavelength of 3 Å, the samples line up perfectly with samples from a 6 Å wave as well. Because we always take the lowest frequency wave that explain our data, we incorrectly interpret our image as resulting from a 6 Å wave. This incorrect result is called “aliasing” and is an important effect in SPA.

The frequency beyond which aliasing occurs is called the Nyquist frequency, and is half the sampling frequency.

$$
f\_{\mathrm{Nyquist}} = \frac{1}{2} \times f\_{\mathrm{sampling}} = 2 \times \frac{1}{\mathrm{pixel\ size}}
$$

Put another way, any spatial features in the object which are smaller than twice the pixel size will be aliased in the image. This results in the commonly quoted maximum resolution for a particle stack of twice the image pixel size.

Frequencies faster than the Nyquist frequency alias to the frequency below Nyquist and the same distance away from it. This operation is commonly described as “folding over” the Nyquist frequency. Note also that in the 3 Å wavelength example above, the aliased 6 Å wave has opposite phase. This is another property of aliasing: the “folding” operation also flips the phase of the aliased wave.

<figure><img src="/files/YAKPNf6JCopBF79oEBA9" alt="An arrow pointing up at a frequency of 0.6 times the sampling frequency is &#x22;folde&#x22; across the Nyquist frequency to become an arrow pointing down (representing a phase shift) at 0.4 times the sampling frequency"><figcaption></figcaption></figure>

{% hint style="info" %}
The phase flipping here is a property of all discrete representations of continuous signals. It is distinct from the specific mechanism of phase inversion that occurs in the electron microscope, which is modeled by the CTF.
{% endhint %}

In the animation below, the moment the true frequency crosses the Nyquist frequency the aliased wave begins to decrease in frequency, rather than increase, since it is folding over the Nyquist frequency.

<figure><img src="/files/0yBqGR1dYbJ944HNz9Wd" alt=""><figcaption></figcaption></figure>

### CTF Aliasing

Just as particle images are represented with pixels in real space, their Fourier transform is represented with pixels in Fourier space. The box in Fourier space is the same size as the box in real space, but each pixel represents a certain range of frequencies rather than a certain region of space. Specifically,

$$
\mathrm{1\ Fourier\ pixel} = \frac{1}{N \times \mathrm{pixel\ size}}
$$

where N is the box size in pixels and the pixel size is the size of a pixel in the real space image. For instance, if a particle is represented with a box of 128 pixels which each represent 2 Å, the Fourier space particle image has a Nyquist frequency of $$\frac{1}{2 \times 128} \mathrm{\AA{}^{-1}} = \frac{1}{256} \mathrm{\AA{}^{-1}}$$. If the CTF oscillates with a frequency greater than this, the contrast transfer function will be aliased.

Consider the true contrast transfer function of an image collected in a typical electron microscope used in single-particle analysis (300 kV, 2.7 Cs) with 2 µm defocus.

<figure><img src="/files/b1adFSefXVmN4eqEor0j" alt="A graph of the contrast transfer function"><figcaption></figcaption></figure>

However, we cannot capture the full, continuous contrast transfer function. We must model it from our images, which are sampled using discrete Fourier pixels. If the Fourier pixels are small, the resulting image of the contrast transfer function may look like this:

<figure><img src="/files/6e56r6pgNZO2IE5Pn6ye" alt="The sane CTF as above, but discretely sampled. A the highest frequencies the amplitude is slightly off, but the overall signal is correct."><figcaption></figcaption></figure>

At the higher frequencies there is some disruption of the amplitude, but the representation is mostly accurate. However, if the Fourier pixels are too large, the contrast transfer function will be significantly aliased.

<figure><img src="/files/35TWBGW4DXwfjYqasipK" alt="The same CTF as above, but now the sampling rate is too low to accurately capture the CTF."><figcaption></figcaption></figure>

In this rather extreme example, the contrast transfer function modeled from the image (black) is significantly different from the true contrast transfer function (grey), especially past approximately 4 Å. This happens because in this region the contrast transfer function starts oscillating more quickly than the Nyquist limit of the Fourier space image. The high-frequency contrast transfer function oscillations are therefore aliased to incorrect lower frequencies.

This effect is more pronounced with:

* higher defocus values, which increase the true oscillation rate of the CTF,
* smaller pixel sizes, which increase the real space Nyquist limit, requiring the Fourier image to represent a larger region of the CTF, and
* smaller box sizes, which reduce the sampling rate in Fourier space.

Since defocus cannot be changed after acquisition, if significant CTF aliasing is observed, the remaining means of removing it are therefore

* downsampling, which “zooms in” on a subregion of the CTF by removing higher frequencies, or
* using a larger box which gives more pixels in Fourier space to represent the CTF.

<figure><img src="/files/F5R6renROkbaYBmFisF5" alt="Four graphs of the CTF are shown. At the top is the continuous CTF. Below that, the aliased CTF shown previously. Third from the top, the aliased CTF is shown for a downsampled image. This produces a CTF which only extends to around 4 Å, where aliasing is not significant. Finally, the CTF is shown from a bigger box. The larger box means the CTF has more samples, reducing aliasing."><figcaption></figcaption></figure>

Of course, the problem of contrast transfer function aliasing is often rendered moot by the poor signal-to-noise ratio of cryo-EM data. For instance, even the severely-aliased example above more-or-less correctly models data up to the resolution to which the CTF fits the data well.


# Symmetry in CryoSPARC

## Symmetry Groups

This page illustrates the various symmetry groups that are supported when enforcing symmetry during a 3D reconstruction in CryoSPARC.

{% hint style="info" %}
Note that simple 3D shapes like prisms are used in the examples below for illustration. These simple shapes also have mirror symmetry, but biological macromolecules are chiral and so cannot be mirrored; only rotations are allowed.
{% endhint %}

### Cyclic (Cn) Symmetry

<div align="center"><figure><img src="/files/WwECrjrPT75Fk4i9GODh" alt="" width="375"><figcaption><p><em>This object is C5 symmetric. A single ASU is highlighted. It can be rotated by 1/5 turn about the Z axis and remain unchanged. It has no other symmetry operations.</em></p></figcaption></figure></div>

The simplest symmetry groups are the cyclic groups, designated Cn. They have the following properties:

* N-fold symmetry about a single axis. By convention in CryoSPARC the symmetry axis is aligned to the Z-axis.
* Symmetry order of N.

{% hint style="info" %}
Note that C1 symmetry (the only symmetry operation is a full 360° rotation) is equivalent to no symmetry.
{% endhint %}

### Dihedral (Dn) Symmetry

<figure><img src="/files/NLlKMNgOHRzbKJyNB8jG" alt="" width="375"><figcaption><p><em>This object is D5 symmetric. A single ASU is highlighted. It can be rotated by 1/5 turn about the Z axis and/or a 1/2 turn about the Y axis and remain unchanged. It has no other symmetry operations.</em></p></figcaption></figure>

Dihedral symmetry groups have the following properties:

* N-fold symmetry about one axis. By convention in CryoSPARC, this axis is aligned to the Z-axis.
* 2-fold symmetry about a second, orthogonal axis. By convention in CryoSPARC, this axis is aligned to the Y-axis.
* Symmetry order of 2N.

### Tetrahedral (T) Symmetry

<figure><img src="/files/hKGgvUTQM8k6wciwlwMz" alt="" width="375"><figcaption><p><em>This object is T symmetric. Note that the base is triangular. A single ASU is highlighted. The 3-fold axes pass through each vertex and the center of the opposite face, while the 2-fold axes pass through the centers of each opposing pair of edges.</em></p></figcaption></figure>

Objects with tetrahedral symmetry have the following properties:

* 3-fold symmetry around four axes. By convention in CryoSPARC, one of these axes is aligned to the Z-axis.
* 2-fold symmetry around three axes. By convention in CryoSPARC, one of these axes lies in the YZ plane.
* Symmetry order of 12.

### Octahedral (O) Symmetry

<figure><img src="/files/aekaLwaqkxbT2kUUQFOB" alt="" width="375"><figcaption><p><em>This object is O symmetric. A single ASU is highlighted. The 4-fold axes are aligned to X, Y, and Z. The 3-fold axes pass through the centers of each opposing pair of faces. The 2-fold axes pass through the centers of each opposing pair of edges.</em></p></figcaption></figure>

Objects with octahedral symmetry have the following properties:

* 4-fold symmetry along three orthogonal axes. By convention in CryoSPARC, the symmetry axes are aligned to the X, Y, and Z axes.
* 3-fold symmetry along four axes.
* 2-fold symmetry along six axes.
* Symmetry order of 24.

### Icosahedral (I) Symmetry

<figure><img src="/files/KAUwR8zNGVr55Mtsfh0Y" alt=""><figcaption><p><em>These objects have icosahedral symmetry. A single ASU is highlighted. The left object is aligned according to the I or I1 conventions, while the object on the right is aligned to the I2 convention. The 5-fold axes pass through opposing pairs of vertices. The 3-fold axes pass through the centers of opposing pairs of faces. The 2-fold axes pass through the centers of opposing pairs of edges.</em></p></figcaption></figure>

Objects with icosahedral symmetry have the following properties:

* 5-fold symmetry along six axes.
* 3-fold symmetry along ten axes.
* 2-fold symmetry along 15 axes.
* Symmetry order of 60. CryoSPARC defines two icosahedral conventions. Both of them align a 2-fold axis to each of the X, Y, and Z axes.
* `I1` (or simply `I`) places the vertices with greatest Z value in the YZ plane.
* `I2` places the vertices with greatest Z value in the XZ plane.


# The FSC and Gold-Standard Refinement

## Overview <a href="#id-2c74324a-3a7b-80be-8692-fe7b9da9157a" id="id-2c74324a-3a7b-80be-8692-fe7b9da9157a"></a>

Single particle analysis attempts to iteratively estimate some unknown quantity (the 3D volume) from a set of incomplete and noisy observations (the 2D particle images). Problems of this nature are potentially susceptible to overfitting, meaning that noise from the data remains or is amplified in the estimated 3D volume.

The Fourier Shell Correlation (FSC) is a measure of correlation between two volumes as a function of signal frequency (Harauz and van Heel, 1986) and can be used in many different ways, for example comparing an experimental 3D density map with a map derived from an atomic model.

If two volumes are independently produced from half of the particle images each, the FSC computed between them is known as the gold-standard FSC (GSFSC) and is one way of measuring the degree of overfitting present at each spatial frequency in a map. The GSCFSC tells us how much we can trust the reconstruction at each resolution since the two volumes will have trustable signal in common, while the noise in each map will be different. The GSFSC's trustability estimate can be used to validate the resulting volume. Building on this idea, gold-standard refinement attempts to limit overfitting from building up iteratively by lowpass filtering the 3D reference volume during each image alignment phase of the [Expectation-Maximization](https://guide.cryosparc.com/expectation-maximization-in-cryo-em) algorithm using the GSFSC curve (Scheres and Chen 2012).

## Before Reading This Page <a href="#id-2a74324a-3a7b-8002-aea1-d49f316c611c" id="id-2a74324a-3a7b-8002-aea1-d49f316c611c"></a>

{% hint style="info" %}
The Fourier transform is a foundational concept for all of signal processing, including cryo-EM. The next section provides a brief overview of the Fourier transform and Fourier space, focusing on the parts which are useful for understanding GS refinement, but it is by no means an exhaustive explanation of the Fourier transform.

If the Fourier transform is completely unfamiliar to you, you may find a more thorough review helpful before proceeding. A good starting place might be [this overview from Better Explained](https://betterexplained.com/articles/an-interactive-guide-to-the-fourier-transform/) or this [article from 3blue1brown](https://www.3blue1brown.com/lessons/fourier-transforms).
{% endhint %}

{% hint style="info" %}
This page assumes some familiarity with the Expectation Maximization algorithm in cryo-EM. If you are not familiar with this topic, we recommend you return to this page after reading the [dedicated Expectation Maximization guide page](https://guide.cryosparc.com/expectation-maximization-in-cryo-em).
{% endhint %}

{% hint style="info" %}
Throughout this page examples are typically given in 2D. This is for convenience of display only — the concepts apply equally well to 3D objects like cryo-EM volumes.
{% endhint %}

## Overfitting <a href="#id-2a74324a-3a7b-80af-a810-e2dc389463a5" id="id-2a74324a-3a7b-80af-a810-e2dc389463a5"></a>

### Overfitting creates spurious features in aligned averages <a href="#id-2b04324a-3a7b-808e-b654-c706ca326398" id="id-2b04324a-3a7b-808e-b654-c706ca326398"></a>

In single particle analysis, we align particle images to a reference and then update the reference using the newly aligned particle images. This iterative process creates a potential for *overfitting*. A volume is *overfit* when it has detail that is purely an artifact of noise in the particle images; over iterations, overfitting can build up as the artefactual details impact subsequent alignments. This problem is often called “Einstein from noise” in homage to [Richard Henderson’s paper on the topic](https://doi.org/10.1073/pnas.1314449110) demonstrating that you can create an image of anything, including Einstein, from pure noise.

We recreate this example here. Consider the following scenario:

* The reference (the object that we align the images to) is a high-resolution picture of Albert Einstein.
* The images are pure Gaussian noise — they have zero real signal of Einstein. However, as with real cryo-EM datasets, we have no way of knowing that the images are junk — we assume they are noisy images of the reference.

When we align all of these images of *random noise* and average them together, we recover an image of Einstein from pure noise.

<figure><img src="/files/wAUuAX6D3cCzthl7dP7H" alt=""><figcaption></figcaption></figure>

This result may be surprising — how could we get an image of Einstein by averaging together thousands of images that do not contain any Einstein signal? This process may be easier to understand with a very simple example. Consider the following setup:

* our reference and particle images have only 25 pixels
* the only allowable poses are rotations by 0, 90, 180, or 270 degrees. No shifts are allowed

<figure><img src="/files/soEceCOtjBqKnk4jmI1K" alt=""><figcaption></figcaption></figure>

The reference has two high-resolution features: single, bright pixels. In the images, on the other hand, some pixels in the noise images are bright and some are dark by pure chance. It is very unlikely that any image, by random chance, has the exact arrangement of dark and light pixels as we see in the reference.

When a noise image is aligned to the reference, the best pose tends to line up the brightest pixels in the image with the reference’s bright pixels. When the noise image is in this pose, the sum of the difference between the image and the reference (the pose's *error*) is minimized. This is all that is meant by calling a pose "best" -- it has the lowest error.

A single image aligned in this way doesn’t look any more like the reference than it did before alignment. There are dark and light pixels scattered throughout the image. However, as we add more and more images, the bright pixels are always aligned to the same place, while the other pixels may be lighter or may be darker. When we average all of these aligned images together, the bright pixels’ consistent placement produces an average which looks like the reference, even though no image looked like the reference on its own.

### Spurious map features from overfitting <a href="#id-2c74324a-3a7b-8046-8aa1-cea0796b0608" id="id-2c74324a-3a7b-8046-8aa1-cea0796b0608"></a>

The images above were created with only a single iteration of alignment to a reference. In a normal 2D Classification or refinement job, several iterations are performed. Because each subsequent iteration uses the previous iteration’s result as a reference, the overfit noise features can compound and become stronger and stronger.

For example, in the first iteration it may be that when all of the noise images were aligned to Einstein, the resulting average looked like Einstein but also had a bright bit of noise somewhere near his face. In the next iteration, this bright stripe might encourage images to put more bright pixels there, making the bright stripe stronger (even though there is no bright stripe in the image themselves or in a true image of Einstein).

This problem can create spurious map features even when the images truly are images of the target object, because *all* images have noise. Overfitting is therefore a critical problem in single particle analysis, but it can be mitigated by only presenting alignment algorithms with a lowpass filtered reference at every iteration.

### Lowpass filtering limits overfitting <a href="#id-2b04324a-3a7b-808a-8426-e3320936e92a" id="id-2b04324a-3a7b-808a-8426-e3320936e92a"></a>

The alignment algorithm was able to create an overfit image because we provided a *high-resolution* reference with fine details. If we lowpass filter the reference before alignment, we still recreate a copy of our input reference. However, because the reference was lowpass filtered, the noise can only be aligned to (and overfit to) low resolution features. Thus, instead of individual locks of Einstein’s hair, we see a light-colored blob:

<figure><img src="/files/nGftghWacPbxzbdubD4u" alt=""><figcaption></figcaption></figure>

We can see why by again considering our simple example.

<figure><img src="/files/MbKJMo3jRS1UnOMSMzou" alt=""><figcaption></figcaption></figure>

The lowpass filtered reference no longer has single bright pixels. Instead, the bright pixels have spread out to make some regions of the reference darker and some lighter. Thus, when we align the same noise image as before, the optimal pose (i.e., the one with the least error) aligns the darkest pixels to these dark regions of the reference. The light pixels are in a different position, but that doesn’t matter since the reference no longer has bright pixels.

In *real images*, the high resolution signal is coupled to the low-resolution signal. If Einstein’s nose (a low-frequency signal) is in the wrong place, the strands of his hair (a high-frequency signal) will be rotated and shifted by the same amount. Aligning one aligns the other. *But pure noise has no such coupling across spatial frequencie*s. The pixel patterns which match Einstein’s hair may be in any position relative to the pixel patterns which match his nose.

{% hint style="info" %}
**Critical point**. For real signal, the best-fitting pose against a low-resolution reference is very similar to the best-fitting pose against a high-resolution reference. For overfit noise, the pose matching a low-resolution reference may be very different from that matching a high-resolution reference.
{% endhint %}

Thus, aligning *real images* to a lowpass filtered reference can recreate a *high resolution* image, because lining up the low-resolution features also lines up the high-resolution features. Consider the result when we align images which are still mostly noise, but have a very small amount of real signal, to the same lowpass filtered reference as in the previous figure:

<figure><img src="/files/3j7W5JofVoMrCyCWcNkT" alt=""><figcaption></figcaption></figure>

The real images were able to recover features that were not present in the reference, while the overfit noise (shown previously) precisely recreated the reference only up to the resolution of signal present in the reference.

This provides a hint as to how we may avoid overfitting in cryo-EM. Any new, *reliable* features which have a resolution finer than the input reference's filter resolution could be real features from images, so we include only those features in the next iteration’s template.

These issues are addressed in gold-standard (GS) refinement (Scheres and Chen 2012) using the Fourier Shell Correlation (FSC) and half maps. The rest of this article covers these critical components of GS refinement in detail, but first we provide a brief background on Fourier space and the Fourier transform.

## Aside: Fourier Space <a href="#id-2a74324a-3a7b-8097-ad4a-ec03c6bff39e" id="id-2a74324a-3a7b-8097-ad4a-ec03c6bff39e"></a>

Consider the image below:

<figure><img src="/files/U0r3377XneAWYeT4nICq" alt="" width="296"><figcaption></figcaption></figure>

You could describe this image by listing the intensity value of every pixel. This is, in fact, what the pixels are. Darker pixels mean the intensity is lower; lighter pixels have a higher intensity. This is *real space*: the value at (0, 0) tells you the intensity of the pixel at (0, 0).

If you needed to quickly describe this image you might instead say, “a wave traveling along the X-axis which oscillates twice”. We can turn this *description of waves* into a coordinate: (2, 0) means the wave oscillates twice in the X direction and 0 times in the Y direction. With these coordinates, we can construct a new image: any pixels with a coordinate which match a wave that’s in the image have a value of 1.0, and other pixels have a value of 0.0.

<figure><img src="/files/IZ3micQrwMFJI1TUS2zU" alt="" width="296"><figcaption></figcaption></figure>

In the image above, a pixel no longer tells us the intensity of the image at a point in space. Instead, a pixel tells us the amplitude of a wave which oscillates in a particular direction and with a particular frequency. This alternate space is called *Fourier space*, and the *Fourier transform* is the function that lets us switch between representing the same signal in real space.

Note that there is actually signal at both (2, 0) and (-2, 0). This is because the image would look the same whether there was a wave moving left-to-right or right-to-left. For our purposes, we can consider Fourier space to be *symmetric about the origin*.

One important note: we have not made the signal any *smaller*. We still have to describe every pixel: Fourier space has zeros everywhere except (2, 0) and (-2, 0). Fourier space represents the same information using the same degrees of freedom, but the *meaning* of a particular point in Fourier space is different from a point in real space.

As the animation below shows, higher frequency signals in real space are further from the origin in Fourier space, and rotating the wave in real space rotates the point in Fourier space.

<figure><img src="/files/PZPZbJYD6rxCL21hrmf0" alt=""><figcaption></figcaption></figure>

### More complex signals are represented with more waves <a href="#id-2a74324a-3a7b-8070-8acc-fade5299da19" id="id-2a74324a-3a7b-8070-8acc-fade5299da19"></a>

Most images are not simple sine waves. However, we can represent complex images as the *sum* of sine waves with different amplitude and phase:

<figure><img src="/files/9pr4zlh1brA0mOxkska0" alt=""><figcaption></figcaption></figure>

Waves may also have different intensities. In the examples above, Fourier pixels are either 0.0 or 1.0, but just as in real space, Fourier pixels can take any value. Larger values indicate that the wave traveling in that direction with that frequency contributes more to the image than Fourier pixels with smaller values.

<figure><img src="/files/rcBztpO8o0JjkD3XGmA4" alt=""><figcaption></figcaption></figure>

In this way, even the most complicated signals can be represented in two ways:

* **Real space:** The value at a position in *real space* is the intensity of the signal at that point in space.
* **Fourier space:** The value at a position in *Fourier space* is the intensity of a wave with a given frequency and direction. Positions further from the center represent waves with higher frequencies.

<figure><img src="/files/vKsjK1iBeB8Mlr3okljM" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
**Critical point**: Fourier theory guarantees that *every* discrete signal, no matter how complicated, can be represented as a sum of a finite number of simple waves.
{% endhint %}

We can now return to GS refinement and the FSC, which takes advantage of the relationship between a signal’s resolution and its distance from the center in Fourier space.

## The Fourier Shell Correlation <a href="#id-2a74324a-3a7b-805f-b90c-c64b853da8e0" id="id-2a74324a-3a7b-805f-b90c-c64b853da8e0"></a>

Imagine we have a perfect map of our target and a reconstruction we made by aligning and averaging noisy images.

{% hint style="info" %}
For now, assume we somehow have access to a perfect map. This makes it easier to motivate and explain the Fourier Shell Correlation. Later we discuss how we can calculate an FSC without access to a perfect map.
{% endhint %}

<figure><img src="/files/MnaEm9gv04lsm4vIDvxJ" alt=""><figcaption></figcaption></figure>

We could then evaluate the quality of our reconstruction by measuring the difference between it and the perfect map (the *error* of our reconstruction). We could do this in real space, by measuring the difference of each pixel:

<figure><img src="/files/qLkRteWyJKzV2BvjBbDB" alt=""><figcaption></figcaption></figure>

Real space error could tell us if certain regions of the reconstruction have higher error than others. In the example above, the pixels at the edges of Einstein’s eyes, tongue, and chin are darker (have more error) than elsewhere. One could imagine this type of analysis telling us that a certain domain of a target protein is worse than the rest of the map in a real cryo-EM sample. Our ultimate goal, though, is to determine whether certain *resolutions*, not *spatial regions*, have more error than others.

If we instead measure the error between each *Fourier pixel* of the Fourier-transformations of the perfect map and our reconstruction, we capture the error of individual waves (combinations of frequency and direction):

<figure><img src="/files/r3KZQcC1xZGwEVHKLmmc" alt=""><figcaption></figcaption></figure>

We see now that the center of Fourier space (lower resolutions) has less error than the edges of Fourier space (higher resolutions). This matches what we see in the real space images — error is evenly spread throughout real space, but features like Einstein’s hair and eyebrows (high resolution features) are more degraded than the overall shape of his head (a low resolution feature).

The Fourier error plot shows the error in individual waves, but we’d like to evaluate an entire frequency at once. Recall that the distance from the origin tells us a wave’s frequency, so all of the waves with a given frequency will be the same distance from the origin. Thus, all of the pixels in a *shell* (a hollow sphere of points) comprise a single resolution in Fourier space.

<figure><img src="/files/KlM3aAzqwazUWjUsFlTU" alt=""><figcaption></figcaption></figure>

Here we take a shell that corresponds to approximately 4 Å. The real space image is not interpretable without the low resolution shells — it is displayed here to emphasize that the one-pixel-wide shell in Fourier space captures information from the entire real-space image.

The plot on the right shows the correlation for this particular Fourier shell. We measure the difference between these two shells using correlation because it always ranges from 0 to 1, while the error depends on the absolute values in Fourier space. We can repeat this process for every resolution shell to calculate the full Fourier Shell Correlation (FSC).

<figure><img src="/files/7yEus6VVu4rrD4Q3jCGt" alt=""><figcaption><p>The Fourier shell currently evaluated. Center: The real space signal corresponding to the current Fourier shell. Right: A plot of each shell’s correlation.</p></figcaption></figure>

In this example, the low-frequency shells are perfectly correlated (i.e., the FSC is 1.0). The real space signals are therefore identical. As the frequency increases (as the shell gets further from the origin), overfit noise accumulates in the reconstruction and the Fourier pixels start to differ. The FSC steadily decreases until it is approximately 0 at the highest resolutions — all of the 2 Å features in this map are noise.

We finally have a measurement of the quality of the reconstruction at each resolution. For resolutions at which the FSC near 1, we can be confident that there is little overfit noise. Resolutions with lower FSC values have more overfit noise, and so are less reliable.

Note that the FSC is a *global* measurement. You can see this in the animation above: when we consider a Fourier shell we are measuring *all of the signal at that frequency* across the entire map. Thus, we might be able to say “Information with a resolution of 3 Å has an FSC of 0.3 for this reconstruction, so is mostly noise”, but we cannot say which specific 2 Å features are real and which are overfit noise.

## Half Sets and the GSFSC <a href="#id-2aa4324a-3a7b-80d7-9da3-cafe99e4c42c" id="id-2aa4324a-3a7b-80d7-9da3-cafe99e4c42c"></a>

The FSC, as described so far, has an obvious problem: you must have a perfect reference to which you compare your reconstruction. If you have a perfect reference there’s no need to perform a reconstruction, or indeed do cryo-EM at all, so this measurement would have limited usefulness. We instead use our data to generate *two* references and compare them to each other.

In so-called “Gold Standard Refinement”, the particles are split into two equal half-sets, often called half sets A and B (Scheres and Chen 2012). Each half set is provided the same starting reference, but after that they are treated totally independently. During each iteration, the particles from half-set A are aligned to their reference and backprojected to produce half-map A. Half-set B produces half-map B in the same way. There are now two *independent* estimates of the true map:

<figure><img src="/files/e4FFnD4sUVyQrhmHuXKV" alt=""><figcaption></figcaption></figure>

Each of these half maps has overfit noise. However, because

* they are refined independently,
* they are images of the same object, and
* the noise is not the same in both half sets

the overfit features differ between the two, but the real image signal is the same. We can then calculate an FSC *between the two half maps*:

<figure><img src="/files/9pUN9o5UYA12ttsSVfOO" alt=""><figcaption></figcaption></figure>

This FSC is often called the Gold-Standard Fourier Shell Correlation, or GSFSC. Gold-standard describes the process of refining half sets independently, while FSC still refers to the Fourier Shell Correlation. The GSFSC does not require a perfect map; we now have a measurement of signal reliability that comes directly from our data!

To reduce overfitting, gold-standard refinement uses the GSFSC as a *filter* on each half map before the next refinement iteration (Scheres and Chen 2012). Since the FSC is near-zero when the signal is mostly overfit noise, that noise will not be present in the next iteration’s reference, and so will not be further amplified.

<figure><img src="/files/df6kPUwwjyOKzA3iqnrH" alt=""><figcaption></figcaption></figure>

## Resolution and the Final Map <a href="#id-2b04324a-3a7b-80c6-9579-ec7c96ee29b6" id="id-2b04324a-3a7b-80c6-9579-ec7c96ee29b6"></a>

<figure><img src="/files/A7nB4aUcp2XHXJzEL6EX" alt=""><figcaption></figcaption></figure>

To review GS refinement so far:

1. Particles are split into independent half sets.
2. These half sets are each aligned to their own reference. In the first iteration this is a shared, lowpass filtered input reference provided by you, but in later iterations the half sets are aligned to their own half-maps.
3. The FSC between the two half maps is calculated. The half maps are then filtered by the FSC. Frequencies with high correlation between the two half maps are preserved, because they likely represent real signal. Frequencies that differ between the two half maps are attenuated, because they contain more overfit noise.
4. The process continues, with each half set aligned to the corresponding filtered half map.

At this point, there are two ingredients missing from a standard GSFSC algorithm. First, we still have no way of determining when a refinement has converged (another way of saying it’s completed). Second, once a refinement has finished we are left with two half maps instead of a single map with information from all the particles.

To solve the first problem, we use a metric called the GSFSC resolution to determine whether the half-maps had improved during an iteration. We iteratively refine the volumes as long as GSFSC improves. Once the GSFSC stops improving, we are left with two half maps which were as good as we could get them. We then solve the second problem by averaging the half-maps together and filtering the *averaged map* by the GSFSC. This produces a single volume which contains reliable information from all of the particles and minimized overfit noise.

### GSFSC resolution, convergence, and the final map <a href="#id-2b14324a-3a7b-80d1-808b-d3f57e6b14bf" id="id-2b14324a-3a7b-80d1-808b-d3f57e6b14bf"></a>

Cryo-EM maps are often published listing a single resolution value. This value is the GSFSC resolution, generally defined as the frequency at which the GSFSC curve first crosses 0.143.

<figure><img src="/files/syTqFkrlzzv4Rn1NucLm" alt=""><figcaption></figcaption></figure>

The 0.143 threshold was proposed as the threshold below which signal cannot be trusted, based on signal-to-noise arguments (Rosenthal and Henderson, 2003). Other thresholds have been proposed (e.g., Rohou 2020), but the field has settled on 0.143 because it is a good estimate of the information content in two noisy half maps.

{% hint style="info" %}
As you will see in later sections of this guide and elsewhere, the GSFSC resolution should always be taken as a suggestion of the *reliability* of a map, *not* as a measurement of what features the map may or may not contain. The GSFSC resolution should be taken as a simple jumping off point for deeper analysis, not the final goal of a single particle analysis project.

Always manually inspect your maps to determine their quality and determine next steps.
{% endhint %}

<details open>

<summary>Why is the resolution cutoff 0.143?</summary>

The argument for this threshold was made in a seminal paper by Rosenthal and Henderson (2003). Here we will cover the basic outline of the justification of 0.143 as the resolution cutoff. The details and mathematics can be found in the original paper and are left to the interested reader.

Note first that, ideally, the GSFSC would tell us something about the signal-to-noise ratio of our reconstruction. That is, we can say something like “If the GSFSC resolution is 3 Å, features in the map which are 3 Å and coarser are at least half real signal”. When the FSC is 0.5, the power in the map is half signal and half noise, so one might select a threshold of 0.5.

However, the GSFSC is comparing two maps which are each made from only half the data, but we care about the signal content of the final, averaged map compared to an ideal map with no noise at all. A threshold of 0.5 for the GSFSC is therefore an underestimate of the information content in the final, averaged map. If we select a threshold of 0.143 for the FSC between two half maps, the correlation between the final, averaged map and the unknown perfect map is 0.5.

</details>

Inspecting the GSFSC over the course of a refinement, we see that the early iterations have a poor GSFSC because they were aligned to a low-resolution reference (in this case, lowpass filtered to the default 30 Å). At each iteration, improved particle poses produce higher-quality half maps, which in turn have better GSFSC curves. This means that the next iteration uses a higher-resolution starting reference, so particle poses improve again.

<figure><img src="/files/v60Rw4gXryiVEQlSYptT" alt=""><figcaption></figcaption></figure>

Eventually, the GSFSC stops improving. This is usually because in late iterations the limit on the map quality are intrinsic to the particle images, rather than being due to an aggressive lowpass filter as in the early iterations. Once the GSFSC resolution stops improving entirely, the refinement is said to have *converged* and is complete.

At this point, the refinement produces the final, whole map by

1. averaging together the two half-maps to take advantage of the full dataset, and
2. filtering the full map by the final GSFSC to hide overfit noise.

## Masking and the GSFSC <a href="#id-2be4324a-3a7b-8086-aa39-d0d1844a7284" id="id-2be4324a-3a7b-8086-aa39-d0d1844a7284"></a>

GSFSC refinement as we’ve describe so far works as follows:

1. Cryo-EM volumes can be represented either in real space (where a voxel represents electron potential at a point in space) or in Fourier space (where a voxel represents the strength of a particular wave).
2. Maps become overfit with noise if a high-resolution reference is provided to noisy images.
3. Lowpass filtering a reference reduces or removes overfitting at high frequencies, but also may limit alignment quality.
4. Producing two independent half maps allows an estimate of which frequencies are overfit and which are true signal, since the half maps will have uncorrelated noise but correlated signal.
5. The Gold-Standard Fourier Shell Correlation (GSFSC) is used to filter frequencies. Each frequency is filtered by how well the frequency correlates between two independent half-maps.
6. Refinements start with a very conservative lowpass filter, typically 30 Å. Each subsequent iteration applies the previous iteration’s GSFSC to the half-maps before alignment. Refinements are considered complete when the GSFSC does not improve between two iterations.
7. The final map is produced by averaging the half maps and filtering by the GSFSC. Averaging the half maps improves signal-to-noise ratio, and filtering by the GSFSC removes noise.

There is one critical ingredient missing from this process: the mask.

Consider these two half maps:

<figure><img src="/files/UbkYfyDYszGPYpync3Y7" alt=""><figcaption></figcaption></figure>

The image content is now centered in the box, with a large region of empty space around it. This more closely models a real cryo-EM dataset, in which particle images are extracted with large boxes to account for [signal delocalization](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/ctf-estimation#the-effects-of-the-ctf-illustrated). Note that the half maps have noise in the *entire image*, even though the real image content is only in the center. This noise is not correlated between the two half maps, so the FSC suffers:

<figure><img src="/files/QBzmDo2EvVG91LWF0vLw" alt=""><figcaption></figcaption></figure>

This is an underestimate of the true reliability of our reconstruction, because it includes noise from regions of space in which there is zero signal. We can instead calculate a *masked* FSC, in which we first mask each half map, then calculate an FSC between these masked half maps.

<figure><img src="/files/T2TrVUKqnXFBc0TfsJ2b" alt=""><figcaption><p>The region outside the mask (i.e., where the mask is set to 0.0) is hatched with white.</p></figcaption></figure>

The masked FSC (solid) is markedly improved from the unmasked FSC (dotted), reflecting the fact that the region of the map we *expect to be correlated* is in fact well correlated. CryoSPARC actually calculates multiple masks and displays curves for each of them. More information about CryoSPARC’s FSC plots is available in [Tutorial: Common CryoSPARC Plots](https://guide.cryosparc.com/processing-data/tutorials-and-case-studies/tutorial-common-cryosparc-plots), and more information about automatic mask generation is available in [3D Masking in Refinement](https://guide.cryosparc.com/processing-data/tutorials-and-case-studies/tutorial-dynamic-masking-in-refinements-v5.0).

### Noise substitution <a href="#id-2c74324a-3a7b-8013-8a13-fc919c5912ed" id="id-2c74324a-3a7b-8013-8a13-fc919c5912ed"></a>

It’s important to note here that the region outside the mask is set to zero, and so correlates perfectly. Thus, if we choose a small, detailed mask we can create an artificially high FSC.

<figure><img src="/files/5rlzz4hBAnxS0kksQp8S" alt=""><figcaption></figcaption></figure>

None of the features visible in the masked map (the two eyes and the smile) are present in the unmasked map. However, because the vast majority of the box is set to 0 by the mask (and the non-zero regions are in exactly the same place), the map has a good FSC curve. Thus, by GSFSC alone, these mask artifacts appear to be reliable map features. To help reduce the effect of masking artifacts on the FSC curve, most cryo-EM refinement software packages calculate *noise-substituted FSC curves* at the end of a refinement (Chen et al. 2013).

First, the phases of every Fourier component in the map beyond a certain resolution. In CryoSPARC, phases are randomized starting with the unmasked GSFSC resolution, or 75% of the masked GSFSC resolution, whichever is coarser. These phase-randomized half maps should have zero correlation in frequency shells with randomized phases.

<figure><img src="/files/Dd9jx19bdslfx9dgg42D" alt=""><figcaption><p>In this example, phases are randomized starting with the unmasked GSFSC resolution (dotted vertical line). Note that, in Fourier space (blue), the maps become random noise starting at the corresponding distance from the origin.</p></figcaption></figure>

Next, the mask is applied to these phase-randomized half maps.The tighter and more detailed the mask, the greater the introduced, artifactual correlation between the two half maps. In this case, because the mask is so small and detailed, the phase-randomized half maps correlate well until almost Nyquist.

<figure><img src="/files/gBjGSw6m9KRKSCoC1sBX" alt=""><figcaption></figcaption></figure>

Now that we know the level of correlation introduced solely by the mask, we can correct the masked GSFSC, attempting to remove this artifactual correlation. Because the low frequencies were not phase randomized, this correction is only performed beginning with the first randomized frequency shell. This produces the corrected GSFSC curve, which ideally tracks the tight mask GSFSC curve. In this case, because essentially all of the high-frequency correlation was induced by our mask, the corrected GSFSC immediately jumps to 0.

<figure><img src="/files/JbcvcnAMwMd2BsWvHhg2" alt=""><figcaption></figcaption></figure>

Compare this corrected curve to that of the loose mask, which in this case is a simple shape surrounding Einstein’s face.

<figure><img src="/files/cGk7ksmcXBruyAiTH0Z0" alt=""><figcaption></figcaption></figure>

Because the mask does not introduce significant bias at the relevant frequencies, the phase-randomized half-maps are essentially uncorrelated. Thus, the corrected curve tracks the masked curve closely, indicating that the GSFSC resolution estimate is not significantly biased by the mask.

{% hint style="warning" %}
CryoSPARC only performs this phase-randomization procedure at the end of the refinement, not during every iteration. The corrected FSC curve therefore indicates *only* whether the mask introduced bias *on its own*. A map with a corrected GSFSC curve that closely tracks the tight mask curve may still have noise that was overfit for other reasons.
{% endhint %}

## GSFSC is a Global Measure <a href="#id-36d4324a-3a7b-805d-a770-c6b2f708d127" id="id-36d4324a-3a7b-805d-a770-c6b2f708d127"></a>

Each Fourier shell considers signal at a given frequency in all directions, across the entire map. This comes with two important considerations.

First, consider a particle stack in which all viewing directions except a select few are well sampled. This produces half maps with high correlation everywhere except a narrow band of each shell in Fourier space, so the overall GSFSC curve remains high even if the map itself becomes unusable. A concrete example of this can be seen in the case of [HA Trimer](https://guide.cryosparc.com/processing-data/tutorials-and-case-studies/case-study-picking-induced-orientation-bias-in-ha-trimer-empiar-10096-and-10097#missing-views-and-anisotropy), and this problem can be addressed with the [Orientation Diagnostics](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/utilities/job-orientation-diagnostics) job.

Second, consider a map in which most of the particle is well-resolved, but a large flexible domain blurs out due to poor alignment. This blurry region has poor correlation between the half maps. Although this poor correlation is localized in *real* space, it is spread throughout all of *Fourier* space, and so hurts the entire GSFSC curve. This affects not only the resolution estimate, but also the filter applied at each iteration of a refinement, potentially limiting the quality of even the well-aligned, stable part of the target. A concrete example of this issue can be seen with [the yeast spliceosome](https://guide.cryosparc.com/processing-data/tutorials-and-case-studies/case-study-yeast-u4-u6.u5-tri-snrnp#global-refinement). The global impact of a poorly-aligned region can be attenuated by either providing a [resolution mask](https://guide.cryosparc.com/processing-data/tutorials-and-case-studies/tutorial-dynamic-masking-in-refinements-v5.0#masking-during-refinement) during the refinement, or by using [Non-Uniform Refinement](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/3d-refinement/job-non-uniform-refinement-new).

## Always Inspect Your Maps <a href="#id-38f4324a-3a7b-80b0-a2b9-ce3a71307631" id="id-38f4324a-3a7b-80b0-a2b9-ce3a71307631"></a>

Although the GSFSC is a useful tool both to prevent overfitting and to provide a summary of a map’s overall quality, it is always essential to evaluate maps by manual inspection. The GSFSC is just one metric, and cannot measure the quality of particular regions of interest, or how two maps may differ in their overall interpretability. Maps should always be downloaded and manually inspected before moving on to downstream analysis steps.

## References <a href="#id-36d4324a-3a7b-80aa-bad2-fe939686c524" id="id-36d4324a-3a7b-80aa-bad2-fe939686c524"></a>

Chen, S. *et al.* High-resolution noise substitution to measure overfitting and validate resolution in 3D structure determination by single particle electron cryomicroscopy. *Ultramicroscopy* **135**, 24–35 (2013).

Harauz, G. & van Heel, M. Exact filters for general geometry three dimensional reconstruction. *Optik.* **73**, 146–156 (1986).

Henderson, R. Avoiding the pitfalls of single particle cryo-electron microscopy: Einstein from noise. *Proceedings of the National Academy of Sciences* **110**, 18037–18041 (2013).

Rohou, A. Fourier shell correlation criteria for local resolution estimation. *bioRxiv* 2020.03.01.972067 (2020) doi:[10.1101/2020.03.01.972067](https://doi.org/10.1101/2020.03.01.972067).

Rosenthal, P. B. & Henderson, R. Optimal Determination of Particle Orientation, Absolute Hand, and Contrast Loss in Single-particle Electron Cryomicroscopy. *Journal of Molecular Biology* **333**, 721–745 (2003).

Scheres, S. H. W. & Chen, S. Prevention of overfitting in cryo-EM structure determination. *Nature Methods* **9**, 853–854 (2012).


# Expectation Maximization in Cryo-EM

An overview of the expectation maximization algorithm.

In statistical estimation, the Expectation Maximization algorithm (hereafter E-M) is a broadly useful method for finding the most likely values of unknown or unobserved quantities given a set of observed data points. E-M is used in a wide range of fields for a variety of purposes, and an in-depth explanation of the theory behind it is beyond the scope of this guide. The goal of this page is to explain the basic concepts of E-M as it pertains to estimation of 2D and 3D reconstructions from cryo-EM data. A basic understanding will make it easier to understand what CryoSPARC jobs are doing, when you might want to change the values of various parameters, and the useful ranges of parameter values.

For cryo-EM data processing, the E-M algorithm is used whenever there is:

1. an unknown, unobserved, but important quantity that we wish to estimate (such as a 2D class average or a 3D density map),
2. a set of observed data that are related to the unknown quantity (such as particle images), and
3. for each observed datapoint, one or more "nuisance" quantities that are unobserved and unknown and are not themselves important (such as the pose of the target molecule in each particle image).

This guide explains E-M in the context of 2D classification, which is the simplest scenario where E-M is applied. The concepts explained herein can also be helpful in understanding how 3D refinement and 3D classification work.

![](/files/p9sd6hHvwlAGS0iaWGnQ)

## Terms

To simplify discussion of E-M, we will first define a few terms.

### Probability and Likelihood

You’re likely already familiar with the concept of probability. For example, if you’re asked how often a fair coin comes up heads or tails, you’d say 50% of the time (because the probability of a fair coin coming up heads is 0.5).

For the purposes of understanding cryo-EM reconstruction, you can consider the terms probability and likelihood to be interchangeable. We will often consider the likelihood of an unknown variable taking on a particular value, *given* the data we have observed. This is known as a conditional likelihood.

### Reference and Pose

{% hint style="info" %}
The **pose** is the rotation and translation necessary to make the **projection** of some **reference** match a collected cryoEM **image**.
{% endhint %}

The goal of most cryo-EM methods is finding the *reference* which produced a given set of particle images. In 2D Classification the *reference* would be the 2D class averages, since we assume that all particles are images of the class averages. In 3D refinements the *reference* would be the 3D volume, since we assume that all particle images are projections of a 3D volume.

{% hint style="info" %}
Note that cryo-EM cannot find the *true volume* which produced a single particle image. Instead, the goal of cryo-EM is to find the *most likely* volume, given the entire set of particle images and our basic assumptions about the cryo-EM image formation process.
{% endhint %}

The combination of rotation and translation which describe the relationship between the reference and the particle image is known as that image’s *pose*. In two dimensions, the pose comprises a single rotation and two translations (X and Y). In three dimensions, there are still only two translations (since the Z translation is essentially modeled by the defocus), but three rotations.

![The 3D pose consists of the three rotations and two translations which make the projection of the reference match a given particle image.](/files/r9bMwCJuKzRbFU1qZPaK)

The 3D pose consists of the three rotations and two translations which make the projection of the reference match a given particle image.

If we knew each particle’s pose with perfect accuracy, we could recreate a 2D class average simply by rotating and translating each particle in that average as described by its pose and averaging them together. Unfortunately, we do not know each particle’s pose. This is the missing, unobserved data. We therefore aim to estimate the most likely class average using our incomplete knowledge of the system, and we will use E-M to do so.

## Expectation and Maximization Steps

{% hint style="info" %}
Particle poses are updated during the **expectation** step. The new poses are used to improve the reference during the **maximization** step.
{% endhint %}

E-M involves the repeated application of two steps: the expectation step and the maximization step.

![](/files/33ASzA06ejnABHAkSRj4)

In the *expectation* step, we use our current estimate of the reference to update the particles’ poses. In the case of 2D Classification, this means comparing the particle image to the current reference and choosing the pose which minimizes the discrepancy between the reference and the image.

In the *maximization* step, we use the new poses found in the expectation step to update the class average. In the case of a 2D Classification, this means taking all of the particle images belonging to a class, rotating and translating them according to their pose, and averaging them together. The maximization step is also often referred to as backprojection or reconstruction.

A single application of these two steps will not find the optimal class average. The poses found during the expectation step likely have some amount of error because they were aligned to a poor initial class average. Likewise, the new class average created by the maximization step was made with imperfect poses and so has some amount of blurring and error. To improve the poses and the class averages, these two steps are iterated. The expectation step is repeated with the new class average, yielding improved pose estimates. These new estimates are used to produce a new class average, which can in turn be used to create a third set of even better pose estimates, etc. This forms the E-M algorithm.

A useful property of the E-M algorithm is that the iterations are guaranteed to converge after a sufficient number of iterations are performed. Moreover, it will converge to the most likely estimate of the unknown variables (the class averages and the poses) if the initial guesses were sufficiently close to correct. There is no universally applicable way to know when the E-M algorithm has converged. In the context of cryo-EM, the steps are often either repeated a pre-defined number of times, or repeated until some metric (like the GSFSC resolution) stops improving.

## Marginalization

{% hint style="info" %}
Marginalizing over a specific variable means accounting for the different values that variable can take and the probability associated with each.
{% endhint %}

During the expectation step, we check how well the image matches the reference from each possible pose. Up to this point, we have been using the single most likely pose as our best guess. Instead, we can assign a likelihood to each pose based on how well the image matches from that pose. Then, during the maximization step, but we can “smear” the particle across every pose, weighted by how likely that pose is. This process of “smearing” is called marginalization.

For a concrete example, consider the following particle image and class average. When the particle image and class average are both noiseless (top) we can assign its pose with a high degree of confidence — it would be surprising if we were wrong. However, when the class average and the particle image are noisy (bottom), we are less confident in our assignment.

![](/files/t7V3Cwz1CEvJ3UE5kcEP)

When we assess poses, we measure the *error* between the pose and the reference (in this case, the reference is the class average). The pose with the notch at the bottom has essentially no error, while the other poses are disagree with the reference to varying degrees. The notch-down pose would be used with a weight of 1.0 in backprojection. More abstractly, we could think of that particle image as contributing all of its “signal mass” to the class average in that pose.

In the noisy case, each of the three poses have some degree of error due to noise in both the image and the reference. Despite the fact that the pose with the notch at the bottom is still the best, we may decide we do not want to discard these other poses which aren’t quite as good, since we cannot be sure they are wrong. We would then *marginalize over pose* by including this particle in the average in each of the three poses, weighted by the likelihood we calculated for that pose. In this case, perhaps we’d assign a weight of 0.75 to the first pose and 0.125 to each of the other two. The particle would then contribute 75% of its “signal mass” to the average in the first pose and 12.5% in each of the other two poses.

*Marginalization* is the general process of allowing a particle to contribute information while accounting for uncertainty in a variable. For 2D classification, for example, those variables are pose and class. The choice to marginalize can be made separately for each variable. For instance, the `Hard classify for last iteration` parameter turns off *marginalization over class* for the final iteration by forcing the particles to only contribute to one class. This case, in which particles only contribute in their best condition, is called *maximization* instead. For a final example, allowing particles to contribute only to their best class, but in all poses weighted by the probability of that pose, would be *maximizing over class, but marginalizing over pose*.

## The Noise Model

The probability of a pose (equivalently, the error between the particle image and the class average) is a function of three things:

1. **The pose itself, i.e., intrinsic error.** There is, ultimately, only one correct pose.
2. **Noise in the particle images.** Noisy images make the correct pose seem worse, since the image is no longer a perfect match for the projected volume. Noisy images may also make wrong poses seem better, if by random chance the noise improves agreement between a wrong pose and the reference.
3. **Noise/errors in the reference.** For example, in the first iteration of 2D Classification, the class averages are initialized essentially randomly. There is not a correct pose in this case — the class averages are fundamentally wrong. On the other hand, in the final iteration of a 2D Classification, the averages are likely quite good, so one particular pose should be better than the rest.

![](/files/RN9izo5X0htGht6Y1Lyx)

Unfortunately, **we have no way of knowing how much of the error in a particular pose is due to each of the three possible sources**, but we only want to discard a pose for having a high **intrinsic** error.

To disentangle the three sources of error, we use a noise model. The noise model tells us how reliable our current system of poses, averages, and images are at each frequency. We can calculate a noise model directly from the images using the current best poses and class averages.

![](/files/IGwWsVsSzW2x7eXt6r4s)

By incorporating the noise model into the error calculation, we become more or less confident in the quality of a pose, depending on the level of noise.

<div data-full-width="true"><img src="/files/LhmggnyDvTr1qsy56vpA" alt=""></div>

## 3D Expectation Maximization

Thus far in this page, we have focused on 2D classification and not 3D refinement. This is solely because it is easier to think about and display 2D operations — the same algorithms and principles apply in the 3D context. For instance, in a Homogeneous or Non-Uniform Refinement, the reference is now a volume and the pose now has three rotations, but the overall iterative process remains the same.

![em-iterations.png](/files/aUc9ZE44dVmNoagghfKs)

Just as in the two-dimensional case, this iterative process improves the quality and accuracy of both the volume and the particles’ pose estimates. Eventually, the process converges — the GSFSC volume stops improving and the particles’ poses stop changing.

![In this animation, each iteration of a Non-Uniform Refinement is shown. The map used for alignment at each iteration is shown on the left. The poses of a random selection of five particles are plotted on the right. In the first iteration, the particles’s poses change by a lot (note the large movement of the yellow triangle). This is because the volume also improves significantly between those two iterations, so the pose estimates become much more accurate. Subsequent iterations fine-tune the poses, but none of the particles move as dramatically after the first.](/files/1WIAPKYbVnHLWAYA9mD9)

In this animation, each iteration of a Non-Uniform Refinement is shown. The map used for alignment at each iteration is shown on the left. The poses of a random selection of five particles are plotted on the right. In the first iteration, the particles’s poses change by a lot (note the large movement of the yellow triangle). This is because the volume also improves significantly between those two iterations, so the pose estimates become much more accurate. Subsequent iterations fine-tune the poses, but none of the particles move as dramatically after the first.

## CryoSPARC Jobs that use Expectation Maximization

<table data-header-hidden><thead><tr><th width="277"></th><th width="397"></th></tr></thead><tbody><tr><td><a href="https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/particle-curation/job-2d-classification">2D Classification</a></td><td>E-M is used to find the most likely pose and class for each particle. Pose marginalization is controlled by <code>Force max over poses/shifts</code>, while class marginalization can be turned off for the final Maximization step by turning on <code>Hard classify for last iteration</code>.</td></tr><tr><td><a href="https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/3d-refinement/job-homogeneous-refinement">Homogeneous Refinement</a> and <a href="https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/3d-refinement/job-non-uniform-refinement-new">Non-Uniform Refinement</a></td><td>E-M is used to find the most likely pose for each particle. Pose marginalization is controlled by the <code>Adaptive Marginalization</code> parameter. <code>Adaptive Marginalization</code> is off by default in Homogeneous Refinement, but on by default in Non-Uniform Refinement.</td></tr><tr><td><a href="https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/3d-refinement/job-heterogeneous-refinement">Heterogeneous Refinement</a></td><td>E-M is used to find the most likely pose and class for each particle. Class marginalization is controlled by the <code>Force hard classification</code> parameter (off by default). Poses are always maximized.</td></tr><tr><td><a href="https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/3d-refinement/job-heterogeneous-reconstruction-only">Heterogeneous Reconstruction Only</a></td><td>Poses and class probabilities are fixed in this job, but class membership is marginalized using existing probabilities unless <code>Force hard classification</code> is turned on.</td></tr><tr><td><a href="https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/variability/job-3d-classification">3D Classification</a></td><td>E-M is used to find the most likely class for each particle (poses are fixed). Class membership is marginalized unless <code>Force hard classification</code> is turned on. <code>Class similarity</code> provides an additional control over how particles are marginalized over class membership.</td></tr><tr><td><a href="https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/local-refinement/job-new-local-refinement-beta">Local Refinement</a></td><td>E-M is used to find the most likely pose for each particle. Poses are marginalized if <code>Marginalization</code> is turned on. Additionally, the likelihood of each pose can be fine-tuned with a prior probability if <code>Use pose/shift gaussian prior during alignment</code> is turned on.</td></tr><tr><td><a href="https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/helical-reconstruction-beta/job-helical-refinement-beta">Helical Refinement</a></td><td>E-M is used to find the pose for each image. No marginalization is performed.</td></tr></tbody></table>


# Get Started with CryoSPARC: Introductory Tutorial (v4.0+)

In this tutorial, we will process a small dataset from movies to reconstructed density map. If you are new to data processing in CryoSPARC, we highly recommend following along!

{% hint style="warning" %}
The information in this section applies to CryoSPARC v4.0+.

For CryoSPARC ≤v3.3, please see: [Get Started with CryoSPARC: Introductory Tutorial (v3)](/guides-for-v3/cryo-em-data-processing-in-cryosparc-introductory-tutorial)
{% endhint %}

## Cryo-EM Data Processing in CryoSPARC: Introductory Tutorial

We recommend starting off with the T20S Tutorial to become familiar with the workflow in CryoSPARC. When you're ready to learn more about cryoEM, you might move on to [our series of introductory videos](/processing-data/tutorial-videos#single-particle-analysis-with-cryosparc) or other [written tutorials](/processing-data/tutorials-and-case-studies).

This dataset is a subset of 20 movies from the [EMPIAR-10025](https://www.ebi.ac.uk/pdbe/emdb/empiar/entry/10025/) T20S Proteasome dataset. While not a representative example of the complexity of most cryo-EM projects today, it is a good way to become familiar with the interface and software features and to [learn how CryoSPARC organizes jobs and projects](/application-guide/projects-workspaces-and-live-sessions).

For a refresher on the interface, projects and jobs, please see the [Application Guide](https://guide.cryosparc.com/application-guide-v4.0+).

![Overview of processing the T20S dataset from raw movie data to a high-resolution 3D structure.](/files/b16TFoFmnMvrUSHAA4v0)

## Introduction: Dashboard, Projects, Workspaces and Jobs

The **Dashboard** provides at-a-glance information on currently active jobs, and your instance's processing history. It also displays links to various resources, including the [CryoSPARC guide](https://guide.cryosparc.com), [Data Processing Tutorials](/processing-data/tutorials-and-case-studies), the [Discussion Forum](https://discuss.cryosparc.com), and recent entries from the [Electron Microscopy Public Image Archive (EMPIAR)](https://www.ebi.ac.uk/empiar/) and [Electron Microscopy Data Bank (EMDB)](https://www.ebi.ac.uk/emdb/). As well, the dashboard displays the [change log](https://cryosparc.com/updates) for new versions of CryoSPARC. The navigation bar on the left side of the UI contains links to the Projects view, the [Resource Manager](/application-guide/managing-jobs), [currently running jobs](/application-guide/managing-jobs), [instance information](/application-guide/instance-management), job history, and more.

CryoSPARC organizes your workflow by **Project**, e.g, P1, P2, etc. Projects contain one or more **Workspaces**, which in turn house **Jobs**.

Projects are strict divisions. Files and jobs from different projects are stored in dedicated project directories and jobs cannot be connected from one project to another.

Workspaces are soft divisions, and allow for logical separation of jobs and workflows so they can be more easily managed in a large project. Jobs may be connected across workspaces and each job may belong to more than one Workspace.

<figure><img src="/files/rxAb49toqjFofL8tGetT" alt=""><figcaption><p>The CryoSPARC dashboard</p></figcaption></figure>

## Step 1: Create a Project

* To create a project, click on the **New Project** button at the top right side of the header.

<figure><img src="/files/NfGPyGyUjKmEiNPK10ql" alt=""><figcaption></figcaption></figure>

* A dialog box will open to the right, prompting you for a project title, container directory, and optionally a description.

<figure><img src="/files/gGJNULI3o7tL20Pn0qKn" alt=""><figcaption></figcaption></figure>

* Enter a project title and browse for a location for the associated container directory with the file browser. The container directory should already exist. CryoSPARC will create a subdirectory within the container directory that will become the **project directory** for the new project, and will act as a root for all new files and directories created in the project. This includes job directories, imported jobs and result groups, and exported jobs and result groups.
* You may also enter a description for your project.
* Click **Create**. The new project now appears on the Projects page, accessible via the container icon on the navigation bar.

## Step 2: Create a Workspace

Use Workspaces to organize or separate portions of the cryo-EM workflow for convenience or experimentation. Create at least one Workspace within a Project before running a Job.

* After creating a new project, you will be prompted to create a new workspace within the project. Enter a title, such as "T20S Subset Processing", and click **Create Workspace**. Note that both the title and description can be modified later.

<figure><img src="/files/FyOZzcIi1ptgm1EnCP82" alt=""><figcaption></figcaption></figure>

* If you have exited the project view, you can always navigate back to it by clicking the container icon on the navigation bar, then clicking anywhere on your newly created project card (in this example, P35), and finally selecting **View Project** on the right-hand sidebar. The container icon and the **View Project** button are highlighted in purple boxes in the image below. Projects can also be navigated to using the search functionality, accessible via the magnifying glass in the lower left corner of the interface:

<figure><img src="/files/P3TNbZCW1zl83AGCKmGH" alt=""><figcaption></figcaption></figure>

* Once within the project, new workspaces can also be created at any time using the **New Workspace** button at the top right side of the header.

## Step 3: Download the Tutorial Dataset

* Log in to the machine where CryoSPARC is installed via command-line.
* Navigate to or create a directory into which to download the test dataset (approx. 8 GB). **This location should have read permissions for** [**the linux user account running CryoSPARC**](https://guide.cryosparc.com/setup-configuration-and-management/cryosparc-installation-prerequisites#2.-common-unix-user-account)**.**
* Run the command `cryosparcm downloadtest` while in this directory. This downloads a subset of the T20S dataset.
* Unpack the downloaded data. The archive file names differ between CryoSPARC versions 4 and 5. In CryoSPARC version 4, run

  ```
  tar xvf empiar_10025_subset.tar
  ```

  In CryoSPARC version 5, run

  ```
  tar xvf empiar_10025_subset_v1.tar 
  ```

## Step 4: Import Movies

* In CryoSPARC, navigate to the new Workspace. To do so, navigate to the project as described in step 2, then click on the workspace card, and hit **View Workspace** on the bottom right.
* You will be greeted with a list of the various [***Import Jobs***](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/import)*,* with links to quickly build any of them. For this tutorial, we can get started with data processing by clicking to build an [***Import Movies***](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/import/job-import-movies) job.

<figure><img src="/files/tVoMlm8vrbv5JmV3xwin" alt=""><figcaption></figcaption></figure>

* This creates a new job within the current Workspace, displayed as a card. By default, new jobs are set to **Building** status, indicated on the job card in purple. To change parameters, select the job and ensure that the job is in Building status. A job's status can be [toggled using the `B` key on your keyboard](https://guide.cryosparc.com/application-guide-v4.0+/creating-and-running-jobs#building-a-job), or by clicking the **Build** or **Stop Building** badges on the job card.

<figure><img src="/files/FHN9govRAiXvVmS5KmBY" alt=""><figcaption></figcaption></figure>

* Select the **Movies data path**: Click the file browse icon and select the movie files. Movies may have a `.tif`, `.eer`, `.mrc` or similar extension. To select multiple files, use an appropriate wildcard expression that matches all desired files. In the case of the specially prepared EMPIAR-10025 movie subset, `14sep05c_*.frames.tif` or simply `*.tif` selects all 20 of the subset's movies. Ensure the wild card expression is both general and specific enough to include all desired an exclude all undesired files. The file browser displays the list of selected files along with the number of matches at the bottom.
* Select the **Gain reference path** with the file browser: In case of the specially prepared EMPIAR-10025 data, select the `norm-amibox05-0.mrc` file in the folder that also contains the `*.tif` movies.
* Edit Job parameters from the Builder; enter the following parameters (obtained from the original publication in [*eLife*](https://elifesciences.org/articles/06380)).
  1. **Raw pixel size (Å)**: `0.6575`
  2. **Accelerating voltage (kV)**: `300`
  3. **Spherical abberation (mm)**: `2.7`
  4. **Total exposure dose (e/Å^2)**: `53`
* After changing a parameter, the parameter box and title changes colour from gray to green. This indicates the parameter is different from its default value:

<figure><img src="/files/qjLMthsPd3HCU2f8D581" alt=""><figcaption><p>Parameters with default values are marked in gray. Custom values are marked in green.</p></figcaption></figure>

* Click **Queue Job** to start the import. Use the subsequent dialog to select a lane/node on which to run the job. The available lanes depend on your installation configuration. By default, import and interactive jobs will run on the master node as they are not resource intensive. Using this dialog box, you may set the job title and description, which may both be changed later. Press the **Queue** button.

<figure><img src="/files/2IZr5Rw1mPkHzRMCJq9E" alt=""><figcaption></figcaption></figure>

* The *Import Movies* job queues and starts running. Look for the Job card in the workspace to monitor its status.
* To open a Job and view its progress, click on the header at the top of the Job card, where the job title is shown. Alternatively, select on the Job card and press the spacebar on your keyboard:

<figure><img src="/files/UQqoFKLi9iWLcVVqHtyU" alt=""><figcaption></figcaption></figure>

* This opens the [Inspect view](https://guide.cryosparc.com/application-guide-v4.0+/inspecting-data#event-log), which shows a streaming event log of the real-time progress for the Import Job. Scroll through the event log to view results. Select a checkpoint to find a specific location in the event log or click 'Show from top' to return to the beginning. Additional actions and detailed information for the job are available in the details panel. The **Output** of the import job, i.e., the 20 imported movies, are available on the right hand side of the event log:

<figure><img src="/files/pgV6QnGiY2BNHoUKs5Pk" alt=""><figcaption></figcaption></figure>

* To exit the job/close the inspect view, press the spacebar again, or press the `×` button on the top-right of the dialog.
* Once finished, the job's status indicator at the top left changes to **"Completed"** in green.

## Step 5: Motion Correction

For our next stage of processing, we will be performing [***Motion Correction***](/processing-data/all-job-types-in-cryosparc/motion-correction) on the imported movies. Motion Correction refers to the alignment and averaging of input movies into single-frame micrographs, for use downstream.

* Select the [**Job Builder**](https://guide.cryosparc.com/application-guide-v4.0+/creating-and-running-jobs) in the right sidebar by clicking on the **Builder** tab. The Job Builder displays all available job types by category (e.g., workflows, imports, motion correction, etc.). A tutorial on the Job Builder and other ways to build jobs in CryoSPARC (Quick Actions, Job Cart) is available [here](https://guide.cryosparc.com/application-guide-v4.0+/creating-and-running-jobs).
* Select the [***Patch Motion Correction***](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction) job type in the Job builder. You can either scroll down and locate the job within the "Motion Correction" category, or you may search for it using the search too&#x6C;*.* This creates a new job in building state so that its inputs and parameters are editable in the right side panel.

<figure><img src="/files/PHOFMF37Vrk6P3AoGilH" alt=""><figcaption></figcaption></figure>

* The *Patch Motion Correction* job requires raw movies as **Inputs**. First, ensure that the *Patch Motion Correction* job is in **Building** status. Open the previously completed **Import Movies** job by clicking on the header of the job card, then drag and drop the **Outputs** of the Import Movies job, to the **Movies** placeholder in the Job Builder.
* Once dropped, the connected output name appears in the Job Builder as an Input:

<figure><img src="/files/VD8U4naRsBmoPxhrgwHr" alt=""><figcaption></figcaption></figure>

* If you have multiple GPUs available, you can speed up the processing time by setting the **Number of GPUs to parallelize** parameter within the **Compute settings** section to the number of GPUs you would like to assign to that job.
* **Queue** the job and select a lane. It is generally not necessary to adjust the *Patch Motion Correction* job parameters; they are automatically tuned based on the data. Finally, click **Queue.**

<figure><img src="/files/SLwUEjv63XViT0nX5Zwr" alt=""><figcaption><p>The job queueing dialog displaying a list of lanes (configured via the CryoSPARC command-line interface) to choose from. Here, we have chosen to queue the job to the <code>cryoem2</code> lane.</p></figcaption></figure>

* Once the job starts to run the card will update with a preview image:

<figure><img src="/files/2AmIYEq1aeM0kSc4yQ2s" alt=""><figcaption><p>The blue icon next to the job ID (J2) indicates this job is running.</p></figcaption></figure>

## Step 6: CTF Estimation

Our next step is to perform [**Contrast Transfer Function (CTF) Estimation**](/processing-data/all-job-types-in-cryosparc/ctf-estimation). This stage involves the estimation of several CTF parameters in the dataset, including the defocus and astigmatism of each micrograph.

* Select [***Patch CTF Estimation***](/processing-data/all-job-types-in-cryosparc/ctf-estimation/job-patch-ctf-estimation) in the Job Builder to create a new job.
* This job type requires **micrographs** as the input. Open the previous *Patch Motion Correction* job, and drag and drop the output (20 micrographs) into the Micrograph placeholder in the Job Builder.

{% hint style="info" %}
You can connect outputs of jobs that haven't completed into the inputs of a building job. In this case, the newly created job will start to run automatically when all parent jobs have completed. This makes it easy to [queue up a series of jobs](https://guide.cryosparc.com/application-guide-v4.0+/creating-and-running-jobs#queuing-chains-of-jobs-that-run-automatically) to run without having to wait until they're completed to queue them manually.
{% endhint %}

* As with *Patch Motion Correction*, the job will complete faster by allocating multiple GPUs. This can be configured with the **Number of GPUs to parallelize** parameter.
* **Queue** the job to start. It is generally not necessary to adjust the *Patch CTF Estimation* job parameters; they are automatically tuned based on the data.

<figure><img src="/files/zVEoRtw6Ho5qL8w9ntQy" alt=""><figcaption><p>Connecting Patch Motion Correction job outputs to a building Patch CTF Estimation job and queuing it to a GPU-accelerated instance.</p></figcaption></figure>

## Step 7: Micrograph Denoiser

{% hint style="info" %}
The Micrograph Denoiser is available in CryoSPARC v4.5 and later. If you are using an older version of CryoSPARC, you can proceed to particle picking without performing Denoising.
{% endhint %}

The motion-corrected micrographs now have CTF estimates. The next step of the Single Particle Analysis pipeline is picking particles, in which each micrograph is scanned for positions which are likely to contain a particle ("picks"). However, cryo-EM micrographs often have very poor signal-to-noise ratio. This makes visual inspection of micrographs difficult, and also makes particle picking difficult, resulting in many off-target picks.

The [Micrograph Denoiser](/processing-data/all-job-types-in-cryosparc/exposure-curation/job-micrograph-denoiser-beta) is a trained neural network that learns to flatten the background noise and increase contrast in particles, making picking more effective. This typically improves performance of all picking techniques, increasing the number of correct particles found and reducing the number of false positives.

* Select *Micrograph Denoiser* in the Job Builder to create a job
* Open the previous *Patch CTF Estimation* job and drag the Micrographs processed output to the exposures input.
* Leave the optional Denoise model input empty, as we have not yet trained a model on this dataset.
* Typically, the parameters of *Micrograph Denoiser* can be left as default. However, since this dataset contains fewer than the default 100 micrographs used for training, we must adjust the **Number of mics for training** parameter to **`20`**.
* **Queue** the job to start.

A denoiser model will be trained on all 20 micrographs of the dataset, and then automatically used to denoise the micrographs for downstream use. More information about how this job works and explanations of the various parameters are available in the [guide page](/processing-data/all-job-types-in-cryosparc/exposure-curation/job-micrograph-denoiser-beta).

<figure><img src="/files/W1PIo17cCcIeV5VTCIHK" alt=""><figcaption><p>Comparison between simple lowpass filtering (left) and denoising (right) of the same micrograph.</p></figcaption></figure>

## Step 8: Particle Picking (Blob Picker)

It is important to attain a large number of high-quality particles for an optimal reconstruction. The [***Blob Picker***](/processing-data/all-job-types-in-cryosparc/particle-picking/job-blob-picker) is a common starting point for particle picking as it is a quick way to obtain an initial set of particle images that can be used to refine picking techniques over time.

Blob picking is a good idea because it verifies data quality and sets expectations for what particle images, projections, and structures should look like. We'll use blob picks to generate a set of templates that can be used as an input to the [***Template Picker***](/processing-data/all-job-types-in-cryosparc/particle-picking/job-template-picker), which will generate a set of much higher-quality picks matching the two primary 2D views of the T20S structure.

{% hint style="info" %}
For certain jobs, CryoSPARC has built-in [**Quick Actions**](https://guide.cryosparc.com/application-guide-v4.0+/creating-and-running-jobs#job-quick-actions). These are shortcuts that allow you to simultaneously build a downstream job while connecting an existing job's outputs to it, all in one step.
{% endhint %}

* To create the *Blob Picker* job using Quick Actions, move your cursor over to the *Micrograph Denoiser* job card and click on the ellipsis (...) in the header of the job card. Alternatively, right click anywhere on the job card.
* Scroll down and click on the **Build Blob Picker on Denoised Mics** button. This will create the *Blob Picker* job and connect the output **Denoised micrographs** from the *Micrograph Denoiser* job to the *Blob Picker* job, all in one click.

<figure><img src="/files/7eD4mHwpoysZ23n0cE1g" alt=""><figcaption><p>Example of building a <em>Blob Picker</em> job from a <em>Micrograph Denoiser</em> job, using Quick Actions.</p></figcaption></figure>

* If you did not use the *Micrograph Denoiser*, select the **Build Blob Picker** option from the Quick Actions of the *Patch CTF Estimation* job instead.
* Once built, enter the Job Builder of the *Blob Picker* job on the right sidebar, and set **Min. Particle Diameter** to **`100`** and **Max Particle Diameter** to **`200`**.
* If you are using denoised micrographs, ensure that **Pick on denoised micrographs** is turned on.

<figure><img src="/files/685iaABtoLTk4gpqA35e" alt=""><figcaption><p>A portion of the <em>Blob Picker</em> event log depicting the template the algorithm uses for picking and exposure image with pick selections plotted as magenta squares.</p></figcaption></figure>

## Step 9: Micrograph Junk Detector

{% hint style="info" %}
The Junk Detector is available in CryoSPARC v4.7 and later. If you are using an older version of CryoSPARC, you can proceed to Inspect Picks using the Blob Picker micrographs and particles.
{% endhint %}

Some contaminants, like carbon support or crystalline ice, are often labeled with many particle picks due to their high contrast. These off-target picks ("junk picks") degrade performance of downstream analysis and should be removed. The Micrograph Junk Detector analyzes the input micrographs and annotates each micrograph with the location of various common types of junk and then removes particle picks from those locations.

* Select *Micrograph Junk Detector* from the Job Builder.
* Open the previous *Blob Picker* job.
* Drag the Micrographs output from *Blob Picker* into the Exposures input of the *Micrograph Junk Detector* builder.
* Drag the All particles output from *Blob Picker* into the Particles input of the *Micrograph Junk Detector* builder.
* Leave all parameters at their default values and launch the job.

The *Micrograph Junk Detector* produces diagnostic plots for the first 20 micrographs (which, for this dataset, is all of them). These plots show which regions of the micrograph have been marked as junk, as well as indicating which particles have been rejected because they are too close to junk.

<figure><img src="/files/PzuuLaS8U0A49Lmc9BuW" alt=""><figcaption></figcaption></figure>

## Step 10: Inspect Picks (Blob Picks)

Use the [***Inspect Particle Picks***](https://guide.cryosparc.com/application-guide-v4.0+/interactive-jobs#interactive-job-inspect-particle-picks)[ ](/processing-data/all-job-types-in-cryosparc/particle-picking/interactive-job-inspect-particle-picks)job to view and interactively adjust the results of blob-based (and template-based) automatic particle picking. Once the previous *Junk Detector* job has completed, create an *Inspect Particle Picks* job by either:

* Using quick actions from the *Junk Detector*, or
* Creating an Inspect Particle Picks job using the builder, then dragging and dropping both the **Particles accepted** and **Labelled micrographs** outputs from the previously completed *Junk Detector* job outputs.

Queue the job.

<figure><img src="/files/8utu8CUESy8z3DHAh9mu" alt=""><figcaption><p>Job details dialog for the <em>Blob Picker</em> open while the <em>Inspect Particle Picks</em> job details panel is active in build mode.</p></figcaption></figure>

* Once the job is ready to interact with, it will be marked as "Waiting" and an "Interactive" tab will be available in the job details dialog:

<figure><img src="/files/GpbCH2uL9ELM453BTLBO" alt=""><figcaption></figcaption></figure>

* The left side of the interactive tab shows three sections. From top to bottom:
  * The **Exposure Plot** displays a customizable scatter plot, allowing various statistics of the exposure dataset (e.g. number of picked particles, average defocus, etc.) to be plotted against each other.
  * Below the **Exposure Plot** is the **Power Histogram**, which displays a 2D histogram of all picked particles. The y-axis measures the **Power Score**, and the x-axis measures the **Normalized Cross-Correlation (NCC)** Score for each particle pick. Note that the histogram includes both true particles, as well as false positives that were picked up during the blob picking. True particles generally have high **NCC** scores (indicating agreement in shape with the templates) and a moderate-to-high **Power** score (indicating the presence of significant signal). Picks that have too little power are false positives containing only ice, while picks with very high power are carbon edges, ice crystals, aggregate particles, etc.
  * Finally, the **Micrographs** tab displays a list of each micrograph in the dataset, along with some statistics for each micrograph. Any micrograph row can be clicked on, to display it on the right panel along with the selected particle picks in green circles.
* Make adjustments to the parameters below if needed. All adjustments are saved automatically. As parameters are adjusted, the selected particle picks on the displayed micrograph will be simultaneously updated.
  * Adjust the Particle Diameter to make it easier to see the location of picks.
  * Clicking the vertical ellipsis button on the top right allows for selection of colours, and whether to display the denoised micrograph or raw micrograph.
  * If viewing raw micrographs, move your cursor over the micrograph display, and adjust the **lowpass filter** slider if needed to better view the picks. It may be easier to view the particles at a lowpass filter value between 20 to 30 Å.

{% hint style="info" %}
The *particle diameter* and *box size* parameters are not used in computing the outputs of this job; Inspect Picks only outputs particle *(x, y)* location coordinates.
{% endhint %}

* Slowly increase the **NCC** slider and watch the particle pick locations on the right. Keep increasing it until empty ice picks stop disappearing and good particles begin to disappear, then reduce it to just below that point. For this example, with denoised micrographs, this value was around **`0.41`**.
* If any picks on empty ice remain, slowly increase the lower **Power threshold** slider until the empty ice picks are removed, but without removing any good particles to do so.
* If any picks remain on high-contrast contaminants, slowly reduce the higher **Power threshold** slider until these junk picks are removed, but without removing any good particles to do so.
* Once satisfied with the picks, select "**Done Picking | Output Locations"** button. This completes the *Inspect Particle Picks* job and saves selected particle locations.

<figure><img src="/files/nwBGRiBdevacCEHOAZqd" alt=""><figcaption><p>Output result groups of the Inspect Picks job showing the resulting 20 micrographs and 12,700 selected particles.</p></figcaption></figure>

## Step 11: Extract from Micrographs (Blob Picks)

This job extracts particles from micrographs, and writes them out to disk for downstream jobs to read.

* Select [***Extract From Micrographs***](/processing-data/all-job-types-in-cryosparc/extraction/job-extract-from-micrographs) in the Job Builder, or by using Quick Actions from the previously completed *Inspect Picks* job.
* Open the recently-completed Inspect Picks job. Drag and drop both the **Micrographs accepted** and **Particles accepted** outputs into the corresponding inputs on the Job Builder.
* In the Job Builder, look under the Particle Extraction section and change the **Extraction box size (pix)** to **`440`**. You may also parallelize the Extract From Micrographs over multiple CPU cores by altering the **Number of CPU cores** parameter, under **Compute settings**.

{% hint style="info" %}
We generally recommend selecting a box size that is at least double the diameter of the particle. The box size controls how much of the micrograph is cropped around each particle location. Larger box sizes capture the most high-resolution signal that is spread out spatially due to the effect of defocus (CTF) in the microscope. However, larger box sizes significantly increase computation expense in further processing. To mitigate this, you can use the [***Downsample Particles*****&#x20;job**](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/particle-picking/job-downsample-particles) to speed up processing for jobs that do not require the full spectrum of data such as 2D Classification.
{% endhint %}

<figure><img src="/files/Dpul4wfUMbegx97Ph5Vl" alt=""><figcaption><p>Workspace with the Extract from Micrographs job builder active.</p></figcaption></figure>

* **Queue** the job.
* Once the job completes, you'll notice the number of resulting particles is less than the input; this is due to the fact that particle extraction process excludes picks that are too close to the edge of the micrograph:

<figure><img src="/files/8CVFhH3oFTlYXdNxPbIU" alt=""><figcaption><p>Output result groups of the Extract from Micrographs job showing the resulting 20 micrographs and 11,129 extracted particles.</p></figcaption></figure>

## Step 12: 2D Classification (To Generate Templates for the Template Picker)

[***2D Classification***](/processing-data/all-job-types-in-cryosparc/particle-curation/job-2d-classification) is a commonly used job in cryo-EM processing to get a first look at the data, group particles by 2D view, remove false positive picks, and even to get early insights into potential heterogeneity present in the dataset. In this step, we will use *2D Classification* to group particles by 2D view, and then use the resulting class averages (also referred to as "templates") to improve our particle picking.

* Select [**2D Classification**](/processing-data/all-job-types-in-cryosparc/particle-curation/job-2d-classification) from the Job Builder, or by using Quick Actions.
* Drag and drop the **Particles extracted** output from the previously completed *Extract from Micrographs*, into the input and queue the job.
* In the event log, preview images of class averages appear after each iteration. Classification into 50 classes (the default number) takes about 15 minutes on a single GPU.

<figure><img src="/files/CKoCV3IcPZbu64WwGIZ8" alt=""><figcaption><p>Generated 2D classes from extracted blob picks.</p></figcaption></figure>

* The quality of classes depends on the quality of input particles. While we could use these blob-picked particles directly for 3D Reconstruction, we can obtain better results by *repeating our particle picking* using these templates generated by 2D Classification. Thus, in this tutorial, we will use these classes to inform the *Template Picker* of the shape of our target structure. The next step in this workflow is the *Select 2D Classes* job.

## Step 13: Select 2D Classes (To Select Templates for the Template Picker)

[***Select 2D Classes***](https://guide.cryosparc.com/application-guide-v4.0+/interactive-jobs#interactive-job-select-2d-classes) allows us to select a subset of the generated templates from *2D Classification*, and to reject the rest.

* Select [***Select 2D Classes***](https://guide.cryosparc.com/application-guide-v4.0+/interactive-jobs#interactive-job-select-2d-classes) from the Job Builder, or by using Quick Actions.
* Drag and drop both the **All particles** and **2D class averages** outputs from the most recently completed *2D Classification* job.
* Queue the job. Once the data is loaded, the job status changes to **Waiting** and Interactive class selection mode is ready.
* Select a "good" class for each distinct view of the structure. In this case, a top view and side view are the most common views present in the dataset. Use both the number of particles and the provided class resolution score to identify good classes of particles. The interactive job provides several ways to sort the classes in ascending or descending order based on:
  * **# of particles**: The total number of particles in each class
  * **Resolution**: The relative resolution of all particles in the class (Å)
  * **ECA**: Effective classes assigned. Classes with a higher ECA value have less confident particle assignments.
* Use the sort and selection controls to quickly sort and filter the class selection. Each class has a right-click context menu that allows for selecting a set of classes above or below a particular criteria.

{% hint style="warning" %}
Avoid selecting classes that contain only a partial particle or a non-particle junk image when creating templates, since these will result in off-center particle picks.
{% endhint %}

<figure><img src="/files/37YfJKoJX6erb0bqm6xm" alt=""><figcaption><p>Interactive Select 2D job depicting two classes selected representing primary orientations.</p></figcaption></figure>

* When finished, select **Done** at the top right side of the window. The job completes.

<figure><img src="/files/H9RRjMDmvcbmKGlYcrd0" alt=""><figcaption><p>Event log of the Select 2D job depicting the two classes selected and 48 classes excluded.</p></figcaption></figure>

## Step 14: Template Picker

The [***Template Picker***](/processing-data/all-job-types-in-cryosparc/particle-picking/job-template-picker) operates similarly to the Blob Picker but allows for an input set of templates to use to more precisely pick particles that match the shape of the target structure.

{% hint style="info" %}
Note that the instructions below apply template picking to the raw micrographs, rather than to the denoised micrographs as we did with blob picking above. In practice, it is recommended to use the denoised micrographs with template picking as well.
{% endhint %}

* Create the *Template Picker* job using the Job Builder.
* Connect the **Templates selected** output of the *Select 2D* into the template input
* Connect the **Micrographs processed** output of the *Patch CTF Estimation* job into the micrographs input
* Set the **Particle diameter (Å)** value to **`190`**
* Queue the job. It should take around 15 seconds to process the dataset.

<figure><img src="/files/v1EkfyYd5MtJBCJfHYQR" alt=""><figcaption><p>Job details dialog depicting the completed Template Picker job with 18,410 picked particles.</p></figcaption></figure>

## Step 15: Inspect Picks (Template Picks)

As with the *Blob Picker*, use the *Inspect picks* job to view and interactively adjust the results of template-based automatic particle picking.

{% hint style="info" %}
Note that you can also use the *Micrograph Junk Detector* job at this stage (as we did above with blob picks), to filter out template picks that are on contaminant regions. The filtered picks can then be used in *Inspect Particle Picks.*
{% endhint %}

* Select ***Inspect Particle Picks*** from the Job Builder, or by using Quick Actions from the previously completed *Template Picker* job.
* Drag and drop both the **All particles** and **micrographs** outputs from the previously completed *Template Picker* job. Queue the job.
* Set the **NCC** and **Power** thresholds using the same process as the previous *Inspect Picks* job. Slowly increase the **NCC** score until good particles start disappearing, then adjust the power scores as necessary.
* Click **Done Picking | Output Locations** to complete the job as before.

<figure><img src="/files/f5rWhYRKYzF16tE4VejQ" alt=""><figcaption><p>Job details dialog depicting the completed Inspect Picks job with 13,808 resulting particles.</p></figcaption></figure>

## Step 16: Extract from Micrographs (Template Picks)

We will now repeat the extraction process given the new set of pick locations generated by the latest *Inspect Picks* job.

* Select [***Extract from Micrographs***](/processing-data/all-job-types-in-cryosparc/extraction/job-extract-from-micrographs)***.***
* Open the recently-completed *Inspect Picks* job. Drag and drop both the **micrographs** and **All particles** outputs into the corresponding inputs on the Job Builder.
* In the Job Builder, look under the Particle Extraction section and change the **Extraction box size (pix)** to **`440`**.

<figure><img src="/files/DOBezS1J2D5G4DJgEy3x" alt=""><figcaption><p>Output result groups of the Extract from Micrographs job showing the resulting 20 micrographs and 12,501 extracted particles.</p></figcaption></figure>

## Step 17: 2D Classification (Template Picks)

* Select **2D Classification** from the Job Builder.
* Drag and drop the **particles extracted** output from the previously completed *Extract from Micrographs*, into the input and queue the job.

<figure><img src="/files/a8RkyWP2tg3Zh47gM706" alt=""><figcaption><p>Generated 2D classes from the template picked particles.</p></figcaption></figure>

* The particles extracted from the template picker result in much higher quality 2D classes. Proceed to the next step to filter out the highest quality classes for 3D reconstruction.

## Step 18: Select 2D Classes

* Select **Select 2D Classes** from the Job Builder.
* Drag and drop both the **All particles** and **2D class averages** outputs from the most recently completed *2D Classification* job.
* Once queued and running, switch to the **Interactive** tab and select all of the good quality classes.
  * In this case, we want to keep all true particles in our dataset (rather than just selecting one top and one side view, as previously done in step 11), and reject all false positives.
  * The visual quality of a 2D class, together with its resolution, number of particles, and ECA, can all provide proxy measurements of the quality of the underlying particles that comprise the class.
  * Note that clear "junk" classes, corresponding to non-particle images, ice crystals, etc., should be rejected at this stage. Err on the side of keeping a 2D class — when a class looks like a blurry version of the particle, it should be kept at this stage. For example, the following shows one possible selection of 2D classes to retain:

<figure><img src="/files/rBmqbw5UKCRhPXMHt1cR" alt=""><figcaption><p>Interactive Select 2D job depicting good quality classes selected.</p></figcaption></figure>

* Once clicking **Done**, the job will generate outputs for each group of classes and particles, one for the selected set and another for the excluded set:

<figure><img src="/files/e2FP9gomfXCPBqBm5Iy5" alt=""><figcaption><p>20 selected classes and 30 excluded classes.</p></figcaption></figure>

## Step 19: Ab-initio Reconstruction

Now that we have a set of good quality particle picks, we can proceed into [*3D Reconstruction*](/processing-data/all-job-types-in-cryosparc/3d-reconstruction)*.*

* Select [***Ab-initio Reconstruction***](/processing-data/all-job-types-in-cryosparc/3d-reconstruction/job-ab-initio-reconstruction) from the Job Builder.
* Drag and drop the **Particles selected** output from the most recently completed Select 2D classes job (classes selected from the result of the template picker) into the **Particle stacks** input in the Job Builder.
* Note: You **do not** need to enforce symmetry during Ab-initio Reconstruction.
* Queue the job. Results appear in real-time in the event log as iterations progress. *Ab-initio* *reconstruction* should generate a 3D density of the T20S structure at a coarse resolution, with no initial model required.

<figure><img src="/files/DfLXbLy2Spl4sRDZSAiu" alt=""><figcaption><p>Event log of the completed <em>Ab-initio Reconstruction</em> job.</p></figcaption></figure>

![Ab-initio volume visualized in UCSF ChimeraX](/files/QVtzAN8H3xMc1pZTRMOb)

## Step 20: Homogeneous Refinement

Now that we have a low-resolution 3D density, we can [*refine*](/processing-data/all-job-types-in-cryosparc/3d-refinement) the density to high-resolution using the [***Homogeneous Refinement***](/processing-data/all-job-types-in-cryosparc/3d-refinement/job-homogeneous-refinement) job.

* Select *Homogeneous Refinement* from the Job Builder.
* Drag and drop both the **All particles** and **Volume class 0** from the recently completed *Ab-initio Reconstruction* into the **Particle Stacks** and **Initial Volume** inputs, respectively. Leave the **Static mask** input empty.
* Set the following parameter:
  * **Symmetry:** `D7`
* Queue the job. Results appear in real time in the stream log. The refinement job performs a rapid gold-standard refinement using the Expectation Maximization and branch-and-bound algorithms. The job displays the current the resolution, measured via [**Fourier Shell Correlation**](https://en.wikipedia.org/wiki/Fourier_shell_correlation), and other diagnostic information for each iteration.

<figure><img src="/files/RBhTSe89s3wl9ISm1qMF" alt=""><figcaption><p>Event log of the completed <em>Homogeneous Refinement</em> job.</p></figcaption></figure>

<figure><img src="/files/yNfQVMiIEr2mqPET29Oq" alt=""><figcaption><p><em>Gold Standard Fourier Shell Correlation</em> (GSFSC) plot for the final refinement iteration.</p></figcaption></figure>

Note that the GSFSC resolution is less than half of the Nyquist resolution for these particles. The particles could therefore be safely downsampled (using [***Downsample Particles***](/processing-data/all-job-types-in-cryosparc/extraction/job-downsample-particles)) to speed up subsequent jobs with no loss in resolution.

Once complete, download the volume and/or mask directly from the **Outputs** section on the right hand side: Select the drop-down to choose the outputs you wish to download. A refinement job outputs a `map_sharp`, the final refined volume with automatic B-factor sharpening applied and filtered to the estimated FSC resolution.

<figure><img src="/files/RblDnBIInXQk41vEicfE" alt=""><figcaption><p>Within the job details dialog, output groups listed have a download menu with various options.</p></figcaption></figure>

![Refined volume visualized in UCSF ChimeraX.](/files/xUdm7EnYd5citkRYiez2)

## Step 21: Mask Generation

{% hint style="info" %}
More information on mask design and creation in CryoSPARC is available in the dedicate [Mask Creation](/processing-data/tutorials-and-case-studies/mask-selection-and-generation-in-ucsf-chimera) tutorial.
{% endhint %}

CryoSPARC automatically generates masks for FSC calculation. However, it is generally good practice to generate your own FSC mask to ensure the calculation is repeatable across jobs, and to ensure the mask covers the entire relevant region of the protein.

First, use *Volume Tools* to lowpass filter the refined volume. This step is important because masks themselves should generally be smooth to avoid introducing bias.

* Create a [***Volume Tools***](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/utilities/job-volume-tools) job and connect the **Refined volume** output from the previous *Homogeneous Refinement* job to the **Input volume** input.
* Set the **Lowpass filter (A)** value to **`10`**.
* Launch the job.

This job lowpass filters the volume, producing a smoothed version suitable for mask creation. Download and inspect the smoothed map in ChimeraX. Find the highest threshold for which no noise is visible in the smoothed map.

<figure><img src="/files/8MKiI8kgCiSGyFmjo6T0" alt=""><figcaption><p>The sharpened map (blue) is displayed with the lowpass filtered map (orange) at various thresholds. The threshold on the left is too high, resulting in too much of the underlying map to be outside the mask. The threshold on the right is too low, resulting in noise in the mask. The center threshold is good, because it preserves all of the information in the map while excluding noise.</p></figcaption></figure>

* Create another *Volume Tools* job and connect the **Volume** output of the previous *Volume Tools* job to the **Input volume** input.
* Set **Type of output volume** to **`mask`**.
* Set **Threshold** to the value you found in ChimeraX.
* Set the **Dilation radius (pix)** to **`4`**. This ensures that all information from the map is contained in the mask.
* Set the **Soft padding width (pix)** to **`14`**. It is important that all masks in CryoSPARC have soft padding added to avoid artifacts.
* Launch the job.

This job creates a mask which can be used in an *Validation (FSC)* job.

* Create a [***Validation (FSC)***](/processing-data/all-job-types-in-cryosparc/post-processing/job-validation-fsc) job from the job builder.
* Connect the **Refined volume** output from Homogeneous Refinement to the **Input volume** input.
* Connect the **Mask** output from the second *Volume Tools* job to the **Static mask** input.
* Set **Compute facility** to **`GPU`**.
* Launch the job.

<figure><img src="/files/6DwAsouzVLaLlKwTbH6o" alt=""><figcaption><p>The event log of a completed Validation (FSC) job.</p></figcaption></figure>

This job produces similar outputs to the GSFSC plots produced by Homogeneous Refinement, but provides better control over the mask used to calculate the GSFSC and allows for repeatable, directly-comparable analysis of resolution.

<figure><img src="/files/W6MvAvG78SdVCFVEJ2Sp" alt=""><figcaption><p>A comparison of the Corrected and Tight FSC curves produced by Homogeneous Refinement (top) and Validation FSC (bottom). In this case, the two curves are very similar. However, the FSC calculated by Homogeneous Refinement has a dip away from the Tight curve, indicating that the mask may have been slightly too tight.</p></figcaption></figure>

## Step 22: Sharpening (Optional)

For optimal results in publications and for model-building, it is often necessary to re-sharpen and adjust the B-factor. This step is optional as Homogeneous Refinement already outputs a sharpened volume (`map_sharp`) with a B-factor reported within the Guinier plot.

* Sharpen the result of the refinement with the [***Sharpening Tools***](/processing-data/all-job-types-in-cryosparc/post-processing/job-sharpening-tools) from the Utilities section in the Job Builder.
* Drag and drop the **volume** output from the result of the previous refinement job, into the **volume** input.
* Set a B-Factor. Get a good starting B-Factor value from the final **Guinier plot** in the stream log of the refinement job you previously ran. It is recommended to input a B-Factor (as a negative value) ±20 the reported value in the plot. In this case, `-56.7` and `-96.7`
* Activate the **Generate new FSC mask** parameter, which will generate a new mask for the purposes of FSC calculation from the input volume. This is done by thresholding, dilating, and padding the input structure.

<figure><img src="/files/y6ERMmLPOL51T9SqEdsg" alt=""><figcaption><p>Guinier plot from the final refinement iteration depicting B-Factor value.</p></figcaption></figure>

* Queue the job. Once complete, download the sharpened map (`map_sharp`) from the output.
* After visually assessing the map, optionally run another sharpening job (clear or clone the existing one) with a different B-Factor.

![Comparison of maps sharpened with different B-factor values. Visualized in UCSF ChimeraX.](/files/UYw4TScPZlrIUN8LTtMj)

## Step 23: Inspect Workflows

Once you have assembled a workflow of connected jobs within a project, you can switch to the [**Tree View**](https://guide.cryosparc.com/processing-data/user-interface-and-usage-guide/job-relationships) to understand how jobs are connected to obtain the final result. Click the flowchart icon in the top header:

<figure><img src="/files/UKb0ER5jmht8UU2tpy5K" alt=""><figcaption><p>Tree view of the full T20S processing workflow.</p></figcaption></figure>

Within the Tree View, you can select jobs and modify/connect them in the same way as previously demonstrated in the Card view. For more information on the Tree view and other useful tips, see the [Application Guide](https://guide.cryosparc.com/application-guide-v4.0+).

## Conclusion

Now that you have refined the data to a high-resolution structure, you can apply more advanced processing techniques. Explore the job builder and other documentation to see the available job types and processing options. Common workflows include:

* [*3D Variability Analysis*](/processing-data/tutorials-and-case-studies/tutorial-3d-variability-analysis-part-one) to explore both discrete and continuous heterogeneity in the dataset
* [*Non-Uniform Refinement*](/processing-data/all-job-types-in-cryosparc/3d-refinement/job-non-uniform-refinement-new) to improve resolutions by accounting for disordered regions and local variations in a structure
* [*Heterogeneous Refinement*](/processing-data/all-job-types-in-cryosparc/3d-refinement/job-heterogeneous-refinement) or [*3D Classification*](/processing-data/all-job-types-in-cryosparc/variability/job-3d-classification-beta) to refine multiple conformations and simultaneously classify particles
  * Sub-classification to identify small, slightly differing populations
* Heterogeneous [*Ab-initio Reconstruction*](/processing-data/all-job-types-in-cryosparc/3d-reconstruction/job-ab-initio-reconstruction) to find multiple unexpected conformational states or multiple distinct particles in the data
* Multiple rounds of [*2D Classification*](/processing-data/all-job-types-in-cryosparc/particle-curation/job-2d-classification) to remove more junk particles
* [*Masked/Local Refinements*](/processing-data/all-job-types-in-cryosparc/local-refinement) to focus on sub-regions of a structure
* Re-pick with multiple higher quality 2D classes
* [*Global*](/processing-data/all-job-types-in-cryosparc/ctf-refinement/job-global-ctf-refinement) (per-exposure-group) or [*Local*](/processing-data/all-job-types-in-cryosparc/ctf-refinement/job-local-ctf-refinement) (per-particle) [*CTF Refinement*](/processing-data/all-job-types-in-cryosparc/ctf-refinement)

For detailed explanations on all available job types and commonly adjusted parameters, see:

{% content-ref url="/pages/-M9xp2LSL4k5GWdrsiL1" %}
[All Job Types in CryoSPARC](/processing-data/all-job-types-in-cryosparc)
{% endcontent-ref %}

{% content-ref url="/pages/-M9yo1eKmRIPG41FztLk" %}
[Tutorial Videos](/processing-data/tutorial-videos)
{% endcontent-ref %}

{% content-ref url="/pages/-MM2IBiuxac1opXNykwH" %}
[Data Processing Tutorials](/processing-data/tutorials-and-case-studies)
{% endcontent-ref %}

Check back to see updates to this guide, as new features and algorithms are in constant development within CryoSPARC.


# Tutorial Videos

CryoSPARC tutorial videos.

## Single Particle Analysis with CryoSPARC

This playlist contains six recordings covering single particle cryo-EM data processing in CryoSPARC from the [2024 Single-Particle Cryo-EM Image Processing Workshop](https://s2c2.slac.stanford.edu/training/workshops-lectures/workshop-lecture-archive) held by the [Stanford-SLAC CryoEM Center (S2C2)](https://s2c2.slac.stanford.edu/training-program-overview) in 2024. The original recordings are courtesy of S2C2.

### Part 1 - Introduction and Cryo-EM Fundamentals

{% embed url="<https://www.youtube.com/playlist?list=PL39mLm0042zJk_ZkM2_OaIoL78uYj6JAb>" %}

In this opening video, the CryoSPARC Team covers fundamental concepts of cryo-EM, including the basics of data collection using an electron microscope, defocus and the contrast transfer function (CTF), motion correction, picking and extraction of individual single particles, expectation maximization and particle classification in 2D and 3D. Be sure to watch the explanation of what a “Think-Pair” question is at 4:45 - they show up throughout the rest of the videos!

In the first section, Data Collection and Preprocessing, we cover what happens to a cryo-EM sample when it’s in the microscope and how this affects the captured images. We also cover the basics of how image aberrations are typically corrected for in cryo-EM.

Next, we discuss how individual particles are picked and extracted from images. If you’ve ever wondered how you should pick a box size, or what the Nyquist Resolution is, this is the section for you!

We then move on to discuss 2D Classification. We also cover the Expectation Maximization algorithm in this section. If this is your first exposure to cryo-EM, or if you’ve always wondered how high-quality 3D maps are actually computed from very noisy images, this is where you can find that information. We also highly recommend all users to watch the section covering the concept of a reference and pose, starting at 40:23. These are critical concepts used in most of the algorithms for single particle cryo-EM analysis.

Finally, we briefly touch on 3D techniques before moving on to discuss validation.

### Part 2 - TRPV1 and a Standard Workflow

{% embed url="<https://www.youtube.com/watch?index=2&list=PL39mLm0042zJk_ZkM2_OaIoL78uYj6JAb&v=MTmv2SIFRb8>" %}

In this video, we present a step-by-step explanation of one possible way to process micrographs of TRPV1 from [EMPIAR 10059](https://www.ebi.ac.uk/empiar/EMPIAR-10059/). The jobs presented here define what we call a “standard workflow”, which we expect to produce a workable consensus refinement for most single particle cryo-EM datasets. By no means is this the only (or the best) way to process data! It is merely meant to guide new users through the most common set of jobs, and present ways to handle challenges as they arise.

Our standard workflow comprises preprocessing, blob picking, particle curation, template picking, more particle curation, and finally a consensus refinement. We cover each of these steps in detail in the relevant chapters, explaining parameter choices along the way. Additionally, we present some jobs which do not produce the desired result and discuss why they failed.

### Part 3 - TRPV5 and Symmetry Breaking

{% embed url="<https://www.youtube.com/watch?index=3&list=PL39mLm0042zJk_ZkM2_OaIoL78uYj6JAb&v=2hBTj1_zWCw>" %}

In this video, we follow the standard workflow established in the previous video to produce a map of a related ion channel, TRPV5 ([EMPIAR 10256](https://www.ebi.ac.uk/empiar/EMPIAR-10256/)). However, each TRPV5 particle binds one calmodulin in one of the four C4 symmetric positions. The C4 symmetric map therefore has artifacts due to misaligned calmodulin molecules.

We compare and contrast three techniques for handling this pseudosymmetry in CryoSPARC. First, we try simply re-refining the particles using a global refinement without imposing C4 symmetry. Next, we try classifying the particles with Heterogeneous Refinement or 3D Classification. Finally, we take advantage of the fact that we know the order of the pseudosymmetry and use Symmetry Relaxation.

After working through the TRPV5 case, we return to TRPV1 and resolve the symmetry mismatch between the C4 symmetric channel and the C2 symmetric DkTx linkers. We try the three methods above again with this new target and compare their performance to the calmodulin case.

Unfortunately, during recording, audio was lost for approximately two minutes from 1:08:05 to 1:10:01. We apologize for the inconvenience.

### Part 4 - Encapsulated Ferritin and Non-Point-Group Symmetry

{% embed url="<https://www.youtube.com/watch?index=4&list=PL39mLm0042zJk_ZkM2_OaIoL78uYj6JAb&v=p_YY5a9apfY>" %}

In this video, we investigate Encapsulated Ferritin (EncFer, [EMPIAR 10716](https://www.ebi.ac.uk/empiar/EMPIAR-10716/)). Each particle in this cryo-EM dataset has two parts: an icosahedral encapsulin and four D5 symmetric EncFer proteins. The EncFer are contained inside the encapsulin protein, meaning there is a symmetry mismatch between the outer shell and the inner parts. The majority of the case study focuses on improving the resolution of the internal EncFer molecules.

Interestingly, although four EncFer molecules appear tetrahedral, their internal D5 symmetry means that the overall arrangement has no point-group symmetry. This requires the use of two new workflows in CryoSPARC, which we call Group Re-alignment and Custom Symmetry Expansion.

A written companion to this video is available in the CryoSPARC Guide: [Case Study: End-to-end processing of encapsulated ferritin](/processing-data/tutorials-and-case-studies/case-study-end-to-end-processing-of-encapsulated-ferritin-empiar-10716).

### Part 5 - FaNaC1 and Discrete Heterogeneity

{% embed url="<https://www.youtube.com/watch?index=5&list=PL39mLm0042zJk_ZkM2_OaIoL78uYj6JAb&v=P5TRDrgDgwk>" %}

In this video, we work on a combined cryo-EM dataset of apo and ligand-bound FaNaC1 (EMPIAR [11631](https://www.ebi.ac.uk/empiar/EMPIAR-11631/) and [11632](https://www.ebi.ac.uk/empiar/EMPIAR-11632/)). Here we present workflows which treat these structures as distinct, discrete states, and discuss important considerations when using jobs like 3D Classification and Heterogeneous Refinement to separate them.

A major point of discussion is the impact of the input poses (from a consensus refinement) on the results of 3D Classification, and how the maps from 3D Classification can be of deceptively poor quality before they are re-refined. During the Q\&A session, we also discuss our opinions on how to decide whether a single particle cryo-EM data processing journey is “complete”, and how this decision depends in large part on the goal of the investigator.

### Part 6 - FaNaC1 and Continuous Heterogeneity

{% embed url="<https://www.youtube.com/watch?index=6&list=PL39mLm0042zJk_ZkM2_OaIoL78uYj6JAb&v=TFcTP33TKUo>" %}

In the final video in this cryo-EM data processing series, we consider the same FaNaC1 dataset as in Part 5, but consider the apo and ligand-bound states as existing on a continuous spectrum. We begin by considering the difficulty inherent in extracting the conformation of individual particle images, and why results from these techniques should always be validated using orthogonal, biochemical methods.

We then discuss the theory and practice of 3D Variability Analysis in CryoSPARC. In 3D Variability Analysis, each particle is modeled as a consensus refinement plus some linear combination of difference volumes. This makes the technique computationally lighter-weight than alternatives and can produce good results in many situations.

Finally, we discuss 3D Flexible Refinement (3DFlex). This technique directly models continuous deformation of the consensus refinement using nonlinear, machine-learning techniques. We discuss important steps in setting up 3DFlex, including mesh design and important parameter choices.

## CryoSPARC Tools: Simple scripting for advanced cryo-EM processing

{% embed url="<https://youtu.be/QY7t67c9bQ8?si=RsBni2ao7WXZSRKp>" %}

CryoSPARC Tools addresses to the need for flexibility in exploring data in intuitive and creative ways, beyond the CryoSPARC™ interface. This open-source Python library lets you script jobs, giving you more control over your data and the ability to do more advanced processing. With it, you can also analyze results, generate plots directly from your metadata, uncover hidden patterns, reproduce images across projects, write data back to CryoSPARC and even integrate other cryo-EM programs. In this video, we demonstrate how CryoSPARC Tools works through a practical example, while also providing an overview of its use cases and applications. Learning resources:

* CryoSPARC Tools website: [https://tools.cryosparc.com/intro.html](https://www.youtube.com/redirect?event=video_description\&redir_token=QUFFLUhqbEIwbWx6UjhJRUdWVWotTUM2LVBFSkZ2SHFPd3xBQ3Jtc0ttY0lhY0FHT2FxaW45TXlvZWJQOUpYUUREVm5YVW9PM09EcVZ0WFVyZTkydkg0b2RtT2cyaENDdDZUZUJVWXY5NVZTazhjM19TcVdGVHhpelc0VmhlM012U1JQUUI1blhYUlBiTnNBMlhnYnRLeVV4MA\&q=https%3A%2F%2Ftools.cryosparc.com%2Fintro.html\&v=QY7t67c9bQ8)
* CryoSPARC examples repository: [https://github.com/cryoem-uoft/cryosp...](https://www.youtube.com/redirect?event=video_description\&redir_token=QUFFLUhqbDBhXzE3c0VCSkY3NUw1RXgxaFNsTWFZRUY1Z3xBQ3Jtc0tuem0tZ2VBQzBZWFRhMDNVZG1GQjZFbWdfV1hqMzlKM0w2Q2I0VC1BQ0tjbFR2VDY0ODdxeUJrM0xRSk94bHV2SjlTOHhsd3ZwaWFMVFJUQzBRTU5SakVLMXlKT0xRc1J4VXEwZFZwRlpKYlNTbGRpVQ\&q=https%3A%2F%2Fgithub.com%2Fcryoem-uoft%2Fcryosparc-examples\&v=QY7t67c9bQ8)
* Discussion Forum (Scripting category): [https://discuss.cryosparc.com/c/scrip...](https://www.youtube.com/redirect?event=video_description\&redir_token=QUFFLUhqbFQ2WnNWN0VlUlJfYS1XekRuUTg3TjZwZ2Jtd3xBQ3Jtc0tua1FiSmRZY0l1SEFqa3J2aHNWV2FZSTBURDJWZW45LXZ2TWh6X0FWcGZOMXZ5Q2tMMmFKM0VkWThXWjlrTlFIOGl3TFFjOENQcjFTQU43S3BhclZWMTVUVVFwZU9pbnVFRmJxSk5nZWZiaXc1dkdtRQ\&q=https%3A%2F%2Fdiscuss.cryosparc.com%2Fc%2Fscripting\&v=QY7t67c9bQ8)
* Example Vesicle Picking: [https://github.com/r-karimi/vesicle-p...](https://www.youtube.com/redirect?event=video_description\&redir_token=QUFFLUhqa0xOcFVrZE1TTU9Qb19IT1Y0TjJNdm9xTXVxQXxBQ3Jtc0tud3M0dzl0ejB2RkJPcW1jU0l4bGtjcEJKZmhhZHZoRml4S1BmX0FwUlNVTXl1R3JCVlJyWnhCaDFsOEtIWkFKQXVFS2JBLVNqR3B4OENmNEZFN2Q0U0gxU3dwR013NWZ2X3VwX21mM3pKRjRWWTdjVQ\&q=https%3A%2F%2Fgithub.com%2Fr-karimi%2Fvesicle-picker\&v=QY7t67c9bQ8)

Familiarize yourself with basic Python knowledge:

* [https://www.python.org/about/gettings...](https://www.youtube.com/redirect?event=video_description\&redir_token=QUFFLUhqbGVPeG9Ra1g0TFVBdFVUZXNaSHVqUnBxempTQXxBQ3Jtc0trenBvNXdYbVA3TTZLaXE0TDItWEpuRmNFOU1vUEU4amxMcVAxNnI1cFV5NUcxcG84c0Jtemp4cXhoRDQ1SGs2RkszRnJtUFpEelVkR1pVaGZxZ2R4YXB3TjZuTlhmb2FqekY5YlF0ZnJwc3RsVTNIVQ\&q=https%3A%2F%2Fwww.python.org%2Fabout%2Fgettingstarted%2F\&v=QY7t67c9bQ8)
* [https://www.codecademy.com/search?que...](https://www.youtube.com/redirect?event=video_description\&redir_token=QUFFLUhqbXJNNVJGc1JFRUtCRkpObk82WUVwWXZGNjluQXxBQ3Jtc0trd2pZRDZ0YXFCMlBLWHpKRncxY1JoNkxCbFNqTS1sRTI3ZHg4clhGM2tSNzIzSXd4Y3dWbEZWdm1ObndGZjJKSEFaZDYwRXZKVnM3c2xTX2NDVEQzSzI1cEd4Y1JpcnN3YW56RnZQS2NUVmVfRVVkcw\&q=https%3A%2F%2Fwww.codecademy.com%2Fsearch%3Fquery%3Dpython\&v=QY7t67c9bQ8)
* [https://diveintopython3.net/](https://www.youtube.com/redirect?event=video_description\&redir_token=QUFFLUhqa0dQcXhlMnUwMHNRX2NZcFdxckowX2FURG5Ud3xBQ3Jtc0tuaDk5aVVSYkZyaGdMdVJLRFJSclN0QjlUdEl5aVJJUjBJSklIbFRoYWlCR0p1RnpTbDNfOXVkMjl4VDVFeWxpVUJWWW1sN2t2aUk5TllFcWVoR1lMeWQzSlNrTDRWQ3NmeWVEYTk5azgzMFVRS2t4aw\&q=https%3A%2F%2Fdiveintopython3.net%2F\&v=QY7t67c9bQ8)
* [https://scipython.com/books/](https://www.youtube.com/redirect?event=video_description\&redir_token=QUFFLUhqa1E0YnVkSkVldTJGZFlhMHZaOExhM0VsaHVrZ3xBQ3Jtc0tuUzZ6NFUySWxabTZoZXkyYWg0RDdzc2xyMXlSWVhBNHZwV19oV2tDWTJabmNyUUVZQldUS05VTmtLTmdsazhHeDEzMjVVQ2pfdFVVV2ZpcGlGOTcxSzJIWVJwelpWVURjaENUV3h0RUtDSEJqXzBjZw\&q=https%3A%2F%2Fscipython.com%2Fbooks%2F\&v=QY7t67c9bQ8)

## Image Processing - 2020 Workshop

{% embed url="<https://www.youtube.com/playlist?list=PL39mLm0042zLabABoJ2zHPR-_3qSPQVWd>" %}

These videos from the 2020 S2C2 workshop cover the v3 interface, but go into more detail about CryoSPARC itself, including how projects, jobs, parameters, and inputs/outputs work in the software. Additionally, these videos investigate two additional datasets: the HA Trimer ([EMPIAR 10097](https://www.ebi.ac.uk/empiar/EMPIAR-10097/)), NaV 1.7 channels ([EMPIAR 10261](https://www.ebi.ac.uk/empiar/EMPIAR-10261/)), and the cannabinoid receptor 1-G protein complex ([EMPIAR 10288](https://www.ebi.ac.uk/empiar/EMPIAR-10288/)).

## CryoSPARC Live Walkthrough

Processing EMPIAR-10288 using CryoSPARC Live.

{% embed url="<https://youtu.be/lrIHBb7Dr8w>" %}

For a detailed guide on CryoSPARC Live, see:

{% content-ref url="/pages/-MNiplu20pJBGs35cP4A" %}
[New Live Session: Start to Finish Guide](/live/new-live-session-start-to-finish-guide)
{% endcontent-ref %}

## Reference Based Motion Correction

How to use Reference Based Motion Correction in CryoSPARC v4.4+.

{% embed url="<https://youtu.be/gnM_IvJShwY?si=IXeoBqiXk4spqke0>" %}

## Workflows

How to construct an automated pipeline of jobs in v4.4+.

{% embed url="<https://youtu.be/Iz7V98aICIo?si=ViW1b4xu92MZ5m89>" %}

## Mask Creation in ChimeraX

Creating a mask in ChimeraX using three different techniques: volume segmentation, volume eraser and molmap.

{% embed url="<https://youtu.be/SIGMDYOC2JM?si=rU6_HsEFxDfWqXwk>" %}

## 3D Flex Custom Mesh Generation

When might your data benefit from a custom mesh, and the process of mesh creation.

{% embed url="<https://youtu.be/KmWyjVSkiic?si=Vp-mWf3XSeQslXKm>" %}


# All Job Types in CryoSPARC

Details on all of the available job types in CryoSPARC, when to use them, required inputs, parameter explanations, and common next steps.

{% content-ref url="/pages/-MUoa87lPm1YW6DAB26E" %}
[Import](/processing-data/all-job-types-in-cryosparc/import)
{% endcontent-ref %}

{% content-ref url="/pages/-MR1pA2tMmOqig-CU8MC" %}
[Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction)
{% endcontent-ref %}

{% content-ref url="/pages/-MR1qLGDMnyAG7KQCg29" %}
[CTF Estimation](/processing-data/all-job-types-in-cryosparc/ctf-estimation)
{% endcontent-ref %}

{% content-ref url="/pages/-MUoamuzjZUCsa1m0-zi" %}
[Exposure Curation](/processing-data/all-job-types-in-cryosparc/exposure-curation)
{% endcontent-ref %}

{% content-ref url="/pages/-MNesJ-dF2JUzMICF7RY" %}
[Particle Picking](/processing-data/all-job-types-in-cryosparc/particle-picking)
{% endcontent-ref %}

{% content-ref url="/pages/f8sIzFKtFclPgNfk7UX6" %}
[Extraction](/processing-data/all-job-types-in-cryosparc/extraction)
{% endcontent-ref %}

{% content-ref url="/pages/-M9xp2LTqFA0SqoyLtJy" %}
[Deep Picking](/processing-data/all-job-types-in-cryosparc/deep-picking)
{% endcontent-ref %}

{% content-ref url="/pages/-MM22IBOzV46KAhZVVtm" %}
[Particle Curation](/processing-data/all-job-types-in-cryosparc/particle-curation)
{% endcontent-ref %}

{% content-ref url="/pages/-MUob\_B0yiYpzeIeiegX" %}
[3D Reconstruction](/processing-data/all-job-types-in-cryosparc/3d-reconstruction)
{% endcontent-ref %}

{% content-ref url="/pages/-MNf7Nb4weQPD3UpCoQP" %}
[3D Refinement](/processing-data/all-job-types-in-cryosparc/3d-refinement)
{% endcontent-ref %}

{% content-ref url="/pages/Cxa5mWaTmQxiDmFoytpO" %}
[CTF Refinement](/processing-data/all-job-types-in-cryosparc/ctf-refinement)
{% endcontent-ref %}

{% content-ref url="/pages/-MUocCsKoGXI9gDlMron" %}
[Conformational Variability](/processing-data/all-job-types-in-cryosparc/variability)
{% endcontent-ref %}

{% content-ref url="/pages/-MUocLY-Uzy34li2rxCk" %}
[Postprocessing](/processing-data/all-job-types-in-cryosparc/post-processing)
{% endcontent-ref %}

{% content-ref url="/pages/-MM26LRRNWHmX2Y0ighO" %}
[Local Refinement](/processing-data/all-job-types-in-cryosparc/local-refinement)
{% endcontent-ref %}

{% content-ref url="/pages/-MNejxB71V\_TUPfVwm8E" %}
[Helical Reconstruction](/processing-data/all-job-types-in-cryosparc/helical-reconstruction-beta)
{% endcontent-ref %}

{% content-ref url="/pages/-MM2DB0zuUBFP8-wJoMk" %}
[Utilities](/processing-data/all-job-types-in-cryosparc/utilities)
{% endcontent-ref %}

{% content-ref url="/pages/IxHDmJ3WEYyXTVKhcFcB" %}
[Simulations](/processing-data/all-job-types-in-cryosparc/simulations)
{% endcontent-ref %}


# Import

Options for bringing data into CryoSPARC for processing.

Import jobs are used to populate a CryoSPARC project with external data sources. Data of all types can be imported.

When large data files (movies, micrographs, particle stacks) are imported into a project, the data files themselves are **not** copied into the project directory, but rather maintained as symlinks to the original files. Therefore, input data for a CryoSPARC project should not be deleted while the project is active.

## Import Jobs

{% content-ref url="/pages/53VGQdSpFPTzXqvX7jBx" %}
[Job: Import Movies](/processing-data/all-job-types-in-cryosparc/import/job-import-movies)
{% endcontent-ref %}

{% content-ref url="/pages/l7oXIP6tBdB9bRbGOvT5" %}
[Job: Import Micrographs](/processing-data/all-job-types-in-cryosparc/import/job-import-micrographs)
{% endcontent-ref %}

{% content-ref url="/pages/wENSi6rB4bUP7jN0vKeh" %}
[Job: Import Particle Stack](/processing-data/all-job-types-in-cryosparc/import/job-import-particle-stack)
{% endcontent-ref %}

{% content-ref url="/pages/K6kV3Z7WbtIUReGq9wJq" %}
[Job: Import 3D Volumes](/processing-data/all-job-types-in-cryosparc/import/job-import-3d-volumes)
{% endcontent-ref %}

{% content-ref url="/pages/sn4miPQZ3fnJXqoWGmc0" %}
[Job: Import Templates](/processing-data/all-job-types-in-cryosparc/import/job-import-templates)
{% endcontent-ref %}

{% content-ref url="/pages/xuGmaeLObE7DVhboIhWX" %}
[Job: Import Result Group](/processing-data/all-job-types-in-cryosparc/import/job-import-result-group)
{% endcontent-ref %}

{% content-ref url="/pages/eQjzpoOJAnwVkNjvj1U8" %}
[Job: Import Beam Shift](/processing-data/all-job-types-in-cryosparc/import/job-import-beam-shift)
{% endcontent-ref %}

## Import Tutorials

{% content-ref url="/pages/-MNf0RAgQryuEQo1LJej" %}
[Tutorial: EER File Support](/processing-data/tutorials-and-case-studies/tutorial-eer-file-support)
{% endcontent-ref %}

{% content-ref url="/pages/-MM2Iwf9GZlJg-2WCA6n" %}
[Tutorial: Negative Stain Data](/processing-data/tutorials-and-case-studies/negative-stain-data)
{% endcontent-ref %}

{% content-ref url="/pages/-MM2JgV53bh37kaK6691" %}
[Tutorial: Phase Plate Data](/processing-data/tutorials-and-case-studies/phase-plate-data)
{% endcontent-ref %}


# Job: Import Movies

## At a Glance

Import one or more raw movies for processing.

## Description

The Import Movies job imports raw Cryo-EM movies into CryoSPARC for end-to-end processing. Importantly, the raw movie files that are imported by this job are *not copied* — they are merely linked (via a symbolic link) into the CryoSPARC project directory where this job is running. It is thus critical that raw movie files are not deleted or moved while they are being processed in CryoSPARC.

{% hint style="info" %}
CryoSPARC is capable of processing raw data in the form of movies and micrographs. Data from modern direct electron detectors typically come in the form of a *movie* (multiple frames), while negative stain data are typically collected as a *micrograph* (one frame). In CryoSPARC, an *exposure* is the generic term referring to the entire data collection event for a single region of the grid and can be either a movie or a micrograph.
{% endhint %}

## Inputs

This job does not accept any inputs.

## **Commonly Adjusted Parameters**

{% hint style="info" %}
If working with negative stain or phase plate data, please see the relevant pages for advice on importing your data to CryoSPARC:

[Tutorial: Negative Stain Data](/processing-data/tutorials-and-case-studies/negative-stain-data)

[Tutorial: Phase Plate Data](/processing-data/tutorials-and-case-studies/phase-plate-data)
{% endhint %}

### Movies data path

The path in which movies are stored. Enter a path, or click on the folder icon to browse or paste the path specifying the location where the movies are stored. To select multiple files, enter a wildcard expression in the browse bar, e.g., `/path/to/files/*.mrc`, which will select all matching file types in the subfolder. Import Movies can accept movies in the `.mrc`, `.mrc.bz2` , `.tif` or `.eer` formats.

### Gain reference path

The path to the gain reference file, if available. Any values in the gain reference which are exactly zero are treated as defects.

{% hint style="info" %}
If you do not have a gain reference image available, first ask your microscope facility for help as it may be stored elsewhere in your data. Otherwise, there are tools which can estimate a gain reference after the fact from your raw data such as RELION’s `relion_estimate_gain`, linked in the [References](#references).
{% endhint %}

<details>

<summary>What is a gain reference?</summary>

Gain references are used to account for differences in the ability of each pixel to detect an electron. A given pixel may be more or less sensitive to electrons than its neighbors. This difference in sensitivity can lead to banding patterns or other artifacts that do not reflect the true sample image. Gain references are special images (typically recorded in a manufacturer-specific format and converted to `.mrc`) which are multiplied by the experimental image to correct for these artifacts.

</details>

### Defect file path

A text file listing defects. Pixels in these regions are filled with noise during motion correction, essentially removing their contribution to the final alignment. This is useful when a row or column of pixels are hot or otherwise incorrect.

#### CryoSPARC defect file format

Each row of a CryoSPARC defect file contains four integers specifying a rectangular defect region. In order, the four integers are:

1. the X position of the left-most edge of the region
2. the Y position of the bottom-most edge of the region
3. the width of the region
4. the height of the region

So, for a micrograph with total width and height of (3710, 3838), a defect file with the lines:

```
1500 0 100 3838
0 500 3710 25
3000 3000 300 50
3000 3000 50 300
```

specifies that the blacked out regions in the image below are defective. Note that in motion correction the defect regions would be filled with noise, not black bars.

<figure><img src="/files/NiyIbY6AoreOIVnA0gtR" alt=""><figcaption></figcaption></figure>

### Microscope parameters

`Raw pixel size` and `Total exposure dose` are typically selected during data collection, while `Accelerating voltage` and `Spherical aberration` are properties of the microscope. These parameters are essential and must be set accurately. If you do not know them, contact the facility at which your data were collected for help.

Data collections are often collected in "super resolution" or "superres" mode. In this mode, electron detection events are localized to "subpixels" which are half the width of the detector's physical pixels. This means that an image is produced with twice as many pixels as the detector actually has; equivalently, the image's pixel size is half that of the detector.

When importing superres movies, you must therefore decide which pixel size to use. When importing movies in the **.tiff** format, you should use the **superres pixel size**. When importing movies in the **.eer** format, you should use the **physical pixel size**.

We recommend Fourier-cropping superres movies back to physical pixel size during Motion Correction, unless it his highly likely that the final map of the sample will have a resolution better than the physical pixel Nyquist (this is relatively rare).

### Negative stain data

If `Negative Stain Data` is on, this indicates that there are light particles on dark background (-1). If it's off, this indicates the movies have dark particles on light background (cryo-em data, +1). Negative stain data are rarely collected in movie format — if you have single-frame micrographs, Import Micrographs is the correct job to use.

### Skip header check

Turned on by default in v4.2+, off by default in prior versions.

When `Skip header check` is turned off, each movie file's header is read by the job to ensure that all movies are of the same size, resolution, and frame count. The header check helps to detect corrupt files which otherwise may cause errors in downstream jobs, but also can take a long time due to the number of file system operations needed to read the headers.

When this parameter is turned off (i.e., when the header check is used), set the `Number of CPUs to parallelize during header check` parameter to parallelize reading of exposure headers.

{% hint style="info" %}
Note that Reference Based Motion Correction and some other jobs require that movies have the same number of frames. [Some users have run into problems](https://discuss.cryosparc.com/t/reference-based-motion-correction-error-all-movies-must-have-the-same-number-of-frames/12740) when using these jobs because the header check was skipped — if you are not certain that all of your movies have the correct number of frames, turning the header check on will avoid these problems down the road.
{% endhint %}

### EER parameters

`EER Number of Fractions` and `EER Upsampling Factor` are only used when importing movies in the [Electron-event representation (EER)](#electron-event-representation-eer) as designed by Guo and colleagues (2020). These parameters determine the number of fractions (roughly equivalent to frames in other formats) and the final resolution sampling of the movie. The defaults are suitable for most cases. For more information, see the EER section below.

{% hint style="info" %}
Note that the default upsampling factor of 2 is equivalent to having collected a super-resolution movie. At this setting, we recommend performing Fourier cropping in subsequent motion correction jobs unless the data are expected to go past the physical Nyquist resolution of the camera.
{% endhint %}

## **Outputs**

### Imported movies

Imported movies are the main expected output and can be used in other CryoSPARC jobs.

### Failed movies

Failed movies are movies that failed the header check and are likely an incorrect size or corrupt. Most often files end up in this output because the gain references or other files were accidentally included in the `Movies data path` wildcard. If the header check is skipped, this output will be empty. It is not necessary to repeat the Import Movies job if you do not need or want to include the movies that failed the header check; you can simply carry on processing with the Imported movies only.

## Common Problems

The Import Movies job will also output thumbnails of the movies with the gain reference applied. It is generally a good idea to check these thumbnails to determine whether flipping or rotation has been applied as expected.

<figure><img src="/files/sKoenguaRhNystLW1MIn" alt=""><figcaption><p>An example thumbnail from an Import Movies job. On the left, the gain reference has not been flipped, yielding thumbnails with large dark and light stripes. Turning on “Flip gain ref &#x26; defect file in Y?” produced the image on the right. Data from EMPIAR 10288 (Kumar et al. 2019).</p></figcaption></figure>

## **Common Next Steps**

Movies must be motion corrected before further processing, typically by CryoSPARC’s [Patch Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction) job.

## Electron-event representation (EER)

Modern direct electron detectors are capable of extraordinarily high frame rates (for example, the Falcon 4 detector has a hardware frame rate of 250 frames per second). However, recording an entire image frame at this frame rate would produce movies with impractically large file sizes. These movies would have frames with almost all pixels dark, almost entirely wasted space. Thus, in typical movie formats (like `.tif` or `.mrc`) a much slower frame rate is used and electron detection events are combined in each of these frames to produce fewer frames with more information per frame.

Instead of recording the entire frame as an image, the EER format records individual electron events by their position and time (Guo et al. 2020). In essence, this allows the movie to record at the full hardware frame rate while producing small movie files. Additionally, the position at which the electron struck the detector can be determined to sub-pixel accuracy, allowing for recording movies at greater than physical resolution, much like Super Resolution modes in other image formats.

Downstream processing still requires traditional images. Thus, the individual electron events are combined when EER files are decoded, into a number of fractions. These fractions contain all of the electron detection events for a given temporal segment of the movie, acting in much the same way as a frame. Higher settings for EER fractions result in potentially higher temporal resolution of sample motion, at the expense of substantially increased processing demands.

Given that EER format records electron events with sub-pixel accuracy, an upsampling factor can be used to decode the files into fractions that are more finely sampled than the physical detector. Guo and colleagues report that the detector is able to capture information at two to three times the physical Nyquist resolution. If a low upsampling factor is used (e.g. 1), the high resolution signal can be aliased to lower frequencies and degrade image quality. They therefore recommend using a high upsampling factor, even as high as 4. If any upsampling is used, we recommend the movies are then Fourier cropped back to physical pixel size (or, for samples expected to achieve Nyquist, a final super-resolution sampling of 2x Nyquist) during motion correction.

## References

1. Guo, Hui, et al. "Electron-event representation data enable efficient cryoEM file storage with full preservation of spatial and temporal resolution." *IUCrJ* 7.5 (2020): 860-869.
2. <https://relion.readthedocs.io/en/release-3.1/Reference/MovieCompression.html#gain-estimation>
3. Kumar, K. *et al.* Structure of a Signaling Cannabinoid Receptor 1-G Protein Complex. *Cell* 176, 448-458.e12 (2019).


# Job: Import Micrographs

## At A Glance

Import one or more micrographs into CryoSPARC for processing.

## Description

The Import Micrographs job imports single-frame micrographs into CryoSPARC for downstream processing. Typically, these micrographs will have been pre-processed through motion correction tools outside of CryoSPARC, or come from a camera that does was imaging single frames rather than movies.

Importantly, the micrograph files that are imported by this job are *not copied* — they are merely linked (via a symbolic link) into the CryoSPARC project directory where this job is running. It is thus critical that the micrograph files are not deleted or moved while they are being processed in CryoSPARC.

The Import Micrographs job accepts an optional input of movies (e.g., from an Import Movies job). This links the micrographs (and therefore, particles extracted from the micrographs) to their source movies, which allows for workflows that require movies such as Reference Based Motion Correction. Thus, although this input is optional, we recommend connecting the movies if they are available.

## **Inputs**

### Source movies (optional)

Imported movies from an Import Movies job. If jobs which require movies (such as [Reference Based Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-reference-based-motion-correction-beta)) will be used later, this input should be connected so that resulting particles are properly linked to their movies. See the [relevant section](#linking-micrographs-to-input-movies) for more information.

## **Commonly Adjusted Parameters**

### Micrographs data path

The absolute path to the micrographs to import (in MRC format). Wildcards can be used to select multiple micrographs.

### Length of prefix to cut

These parameters are used to properly link micrographs to their input movies. See the [Linking Micrographs to Input Movies](#linking-micrographs-to-input-movies) section for help setting these parameters.

### Microscope parameters

`Raw pixel size` and `Total exposure dose` are typically selected during data collection, while `Accelerating voltage` and `Spherical aberration` are properties of the microscope. It is best to specify these properties if they are known. Note that the pixel size of the micrographs may not be the physical detector pixel size if the original movies were downsampled during motion correction or if they were collected in a super-resolution mode.

### Negative stain data

If `Negative Stain Data` is on, this indicates that there are light particles on dark background (-1). If it's off, this indicates the micrographs have dark particles on light background (cryo-em data, +1).

### Skip header check

Turned on by default in v4.2+, off by default in prior versions.

When `Skip header check` is turned off, each micrograph file's header is read by the job to ensure that all micrographs are of the same size in pixels and physical extent. The header check helps to detect corrupt files which otherwise may cause errors in downstream jobs, but also can take a long time due to the number of file system operations needed to read the headers.

When this parameter is turned off (i.e., when the header check is used), set the `Number of CPUs to parallelize during header check` parameter to parallelize reading of exposure headers.

## **Outputs**

### Imported micrographs

The imported micrographs are ready for further processing.

### Failed micrographs

If `Skip Header Check` is off, micrographs with an inconsistent size are output here and are typically not used in downstream processing.

## **Common Next Steps**

Imported micrographs are typically taken into a [Patch CTF](/processing-data/all-job-types-in-cryosparc/ctf-estimation/job-patch-ctf-estimation) job to estimate their contrast transfer function parameters.

## Linking Micrographs to Input Movies

In the scenario where you want to link your micrographs to a movie dataset that is already in CryoSPARC, you can connect the movies to the inputs of this job.

<figure><img src="/files/73K4y2knHIvSVMHF1C3h" alt=""><figcaption></figcaption></figure>

Since CryoSPARC does not know how the micrograph names correspond to movie names, you will have to help the job connect the correct movie and micrograph. Most motion correction packages add characters to the beginning and/or end of a filename, but leave some identifying region in the middle of the filename unchanged. The Import Micrographs job thus provides four parameters, each of which accepts a number telling the job how many characters to trim from the beginning and end of the micrograph and movie filenames to be left with this identifying region.

Consider the following case:

```
Importing movies from /bulk8/data/rposert-dev/CS-rposert-guide-work/J4/motioncorrected/*.mrc
Importing 24 files
Attempting to find corresponding filenames in imported micrographs and connected input movies..
  Example source (input) movie filename:
 003573716055357609725_CB1__00006_Feb18_23.39.55.tif
  Example query (imported) micrograph filename:
 001334752721923690937_CB1__01978_Feb20_11.29.13_background.mrc
```

In these lines from the Import Micrographs job log, we see that the movies have a unique ID (twenty-one numbers) added to the front, and they end in `.tif`. The micrographs also have a (different) twenty-one character ID, but end in `_background.mrc`. We therefore need to cut twenty-one characters off the front of both the movies and the micrographs, four characters off the end of the movies, and fifteen characters off the end of the movies.

<figure><img src="/files/QQuVxh72hTyHw0O9Sxh7" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
The Import Micrographs job displays the first filename in *alphabetical order* in the job log. Since many programs add a unique ID to the beginning of the filename during import and motion correction, the filenames displayed in the log will likely not have any part of their filename in common. You are looking for the part of the filename which *would match* for a movie derived from an imported micrograph.
{% endhint %}

The parameters for the job should therefore be:

```
Length of movie path prefix to cut : 21
Length of movie path suffix to cut : 4
Length of mic. path prefix to cut  : 21
Length of mic. path suffix to cut  : 15
```

Now the imported micrographs will be linked to the movies in CryoSPARC, allowing you to come back to preprocessing steps easily.


# Job: Import Particle Stack

## At A Glance

Import a stack of particles with metadata and CTF parameters.

## Description

The Import Particles job import a stack of extracted particle images (each being a square cropped image of one particle) into CryoSPARC for downstream processing. The particle stack must be in `.mrc` or `.mrcs` format, and the metadata for the particles in `.star` format.

Importantly, the particle stack files that are imported by this job are *not copied* — they are merely linked (via a symbolic link) into the CryoSPARC project directory where this job is running. It is thus critical that the particle stack files are not deleted or moved while they are being processed in CryoSPARC.

## **Inputs**

### Source Exposures (optional)

Some jobs require knowing where in a movie a particle image comes from. For instance, [Remove Duplicate Particles](/processing-data/all-job-types-in-cryosparc/utilities/job-remove-duplicate-particles) requires knowing how close to images are on the movie, and [Reference Based Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-reference-based-motion-correction-beta) corrects motion in each frame of the input movie for each particle. To enable these jobs with imported particles, you must link them to exposures imported separately (e.g., through an [Import Movies](/processing-data/all-job-types-in-cryosparc/import/job-import-movies) or [Import Micrographs](/processing-data/all-job-types-in-cryosparc/import/job-import-micrographs) job).

If Source Exposures are provided as input, the job will attempt to link the imported particles to exposures from which they were extract. In this case, the filenames must match those found in the `rlnMicrographName` of the `.star` files found in `Particle meta path`. Parameters to trim prefix and suffixes of the input exposure names and the names found in `rlnMicrographName` are provided. See [Import Micrographs](/processing-data/all-job-types-in-cryosparc/import/job-import-micrographs#linking-micrographs-to-input-movies) for more information on the use of this type of parameter.

## **Commonly Adjusted Parameters**

### **Ignore raw data**

If this parameter is turned on, particle images will not be imported. The output particles will have to be extracted from micrographs before they can be used in downstream jobs. This is useful when you only want to import particle *locations* from an external source.

### Ignore pose data

If this parameter is turned on, 2D and 3D pose data is discarded from the imported particles.

### Particle meta path

The `Particle meta path` is required, and contains information about the particles, such as their defocus or pose. This is typically some form of text file, for instance a RELION `.star` file.

### Particle data path

The `Particle data path` is required, and contains the particle images themselves.

### Microscope parameter overrides

One can specify `Pixel size (A)`, `Accelerating voltage (kV)`, `Spherical aberration (mm)` and `Total exposure dose (e/A^2)` if known. Otherwise, they will be read from the files found in `Particle meta path`.

`Data Sign` determines whether the information in the particle stack is dark particles on a light background or light particles on a dark background. CryoEM data is typically recorded dark-on-light (`Data Sign = +1`) but can be flipped during processing or extraction in some programs. For instance, RELION flips particles to be light on dark by default during extraction.

CryoSPARC always displays particles with the Data Sign applied; if particles images are stored on disk as light-on-dark and with `Data Sign` set to `-1`, the particles will appear dark-on-light.

## **Outputs**

### **Imported Particles**

An imported particle stack for further use in CryoSPARC.

## **Common Next Steps**

Depending on the source of the particles, it may be appropriate to proceed with particle curation (e.g., [2D Classification](/processing-data/all-job-types-in-cryosparc/particle-curation/job-2d-classification) or [Ab-Initio Reconstruction](/processing-data/all-job-types-in-cryosparc/3d-reconstruction/job-ab-initio-reconstruction)) or other types of refinement (e.g., [Non-Uniform Refinement](/processing-data/all-job-types-in-cryosparc/3d-refinement/job-non-uniform-refinement-new), [3D Flexible Refinement](/processing-data/all-job-types-in-cryosparc/variability/job-3d-flexible-refinement-3dflex-beta)).

## Common Problems

### Incorrect defocus due to higher-order CTF parameters

RELION and CryoSPARC record higher-order CTF parameters in mutually incompatible formats. Performing higher-order refinements in global CTF in one processing software will require re-fitting CTF in the other.

This is because higher-order effects in the CTF account for some of the aberrations that are otherwise incorporated into the defocus. Since RELION cannot read CryoSPARC’s values for these higher-order effects (and vice-versa), the fitted CTF is now incorrect because it has the modified defocus value but not the higher-order effects. This problem is resolved by re-refining the CTF whenever particles with higher-order aberrations are moved between the two packages.


# Job: Import 3D Volumes

## At a Glance

Import one or more volumes from the filesystem or EMDB.

## Description

The Import 3D Volumes job type allows you to import one or more 3D volumes, e.g., half-maps, sharpened or unsharpened maps, local resolution maps, and masks. In CryoSPARC v3.3+, you also have the option to download volumes directly from [EMDB](https://www.ebi.ac.uk/emdb/).

## **Inputs**

This job does not accept any inputs.

## Commonly Adjusted Parameters

### Volume data path

A valid path referencing one or more volumes on the filesystem. Multiple volumes can be specified using wildcards (`*`) and other Unix-style pathname patterns.

### EMDB ID

A comma-separated list of EMDB identifiers (the 4 or 5 digits following "EMD-").

If one or more valid EMDB ID is specified, the job will attempt to connect to EMDB servers and download the entry metadata and volume data. The JSON metadata associated with the entry will be stored in the job directory and the 3D volume will be stored as a `.mrc` for downstream processing.

{% hint style="info" %}
The CryoSPARC master instance must allow outbound requests to the EMDB REST API (HTTPS) at `https://www.ebi.ac.uk/pdbe/api/emdb/all/*` and the EMDB FTP server at `ftp://ftp.ebi.ac.uk/pub/databases/emdb/structures/*` for this functionality to work.
{% endhint %}

### Type of volume being imported

This job accepts maps, half-maps, sharpened maps, local resolution maps, or masks. All of these volumes are ultimately 3D arrays of voxels, but CryoSPARC imposes limits on which volumes can be plugged into which inputs to avoid, e.g., accidentally trying to refine particles against a binarized mask. Selecting the correct value from this dropdown allows these safeguards to work correctly.

### Pixel size (A)

If it is not specified here, the pixel size will be read from the map file header.

## **Outputs**

Imported 3D volume(s) of the appropriate type are produced in separate outputs.

## Common Next Steps

Imported mask bases can be thresholded, dilated, and padded with [Volume Tools](/processing-data/all-job-types-in-cryosparc/utilities/job-volume-tools).

Imported maps can be used in a wide variety of jobs, including [Homogeneous](/processing-data/all-job-types-in-cryosparc/3d-refinement/job-homogeneous-refinement) or [Heterogeneous Refinement](/processing-data/all-job-types-in-cryosparc/3d-refinement/job-heterogeneous-refinement).

## Common Problems

Selecting the incorrect volume type (i.e., `map` when the volume is actually a `mask`) will prevent the volume’s use in the desired slot in downstream jobs. Re-importing the volume with the correct type will resolve the issue.


# Job: Import Templates

## At a Glance

Import one or more 2D template images for particle picking.

## **Inputs**

This job does not accept any inputs.

## **Commonly Adjusted Parameters**

### Templates data path

An absolute path to the `.mrc` file containing the 2D templates.

### Pixel size (A)

If this parameter is not set, the pixel size will be read from the `.mrc` header.

## **Outputs**

### Imported templates

This output contains the templates for further use.

## **Common Next Steps**

The imported templates should only be used for [Template Picking](/processing-data/all-job-types-in-cryosparc/particle-picking/job-template-picker).


# Job: Import Result Group

## At a Glance

Import a Result Group from another job or project.

## Description

CryoSPARC jobs produce Result Groups, which are containers for outputs and associated metadata. These contain a variety of objects, depending on the job type. For instance, a Particles Result Group from a Non-Uniform Refinement contains information about the particles (poses, CTF fits, source micrographs, extracted particle stack locations, etc.). The Volume Result Group contains several volumes (the final reconstruction and the sharpened map, both half maps, and various masks).

These Result Groups can be exported in the Outputs screen. They can then be imported using this job (Import Result Group) to allow for their use in other projects or instances of CryoSPARC.

<figure><img src="/files/YhherUCjR4TTizWsO8vj" alt=""><figcaption><p>The export button can be found on the left-hand side of each result group’s box in the Outputs tab of a completed job.</p></figcaption></figure>

As an example, say project P12 is the first time data has been collected on a certain target. It only went to moderate resolution after processing, so a new sample was produced with improved sample preparation conditions. This data was imported into a new project P13. A user could export the volume group from a refinement in P12, copy the group into P13’s project directory, and import it into P13 to generate templates, skipping the steps of blob picking and template generation in P13.

## Inputs

This job does not accept any inputs.

## Commonly adjusted parameters

This job requires an **absolute path** to the result’s `.csg` file. Absolute paths start with a slash (`/`).

{% hint style="info" %}
If you are viewing the `.csg` file in a terminal window, the absolute path is given by `realpath {filename.csg}`.
{% endhint %}

## Outputs

The outputs of this job depend on the Result Group that is imported.


# Job: Import Beam Shift

## **At A Glance**

Add beam shifts to existing exposures.

## **Description**

The Import Beam Shift job is a new job as of CryoSPARC v4.4, created to add EPU session beam shift information to existing exposures datasets in CryoSPARC, without need for re-importing the movies/micrographs from scratch. For new movie and micrograph imports, beam shift info can be directly imported in the [Import Movies](/processing-data/all-job-types-in-cryosparc/import/job-import-movies) and [Import Micrographs](/processing-data/all-job-types-in-cryosparc/import/job-import-micrographs) jobs.

{% hint style="info" %}
For additional information on importing movies with XML files, refer to the [EPU AFIS Beam Shift Tutorial](https://guide.cryosparc.com/processing-data/tutorials-and-case-studies/tutorial-epu-afis-beam-shift-import).
{% endhint %}

## **Inputs**

### Exposures

An existing exposure (movie and/or micrograph) dataset in CryoSPARC.

## **Commonly Adjusted Parameters**

### EPU XML metadata path

An absolute path, wildcard-expression (e.g. `/mount/data/somewhere/*.xml`) for importing EPU XML files with beam shift. Only `.xml` files are supported.

### Path matching parameters

These four parameters exist to assist in matching movie/micrograph paths (stored in `movie_blob/path` or `micrograph_blob/path`, respectively) to the XML file paths imported from the above `EPU XML metadata path` wildcard expression. Examples of the trimmed file paths will be printed to the event log to help determine the number of characters. The values of these parameters is most quickly determined by running this job with all defaults, and observing the event log. See the [Import Micrographs](/processing-data/all-job-types-in-cryosparc/import/job-import-micrographs#linking-micrographs-to-input-movies) page for more information on how to use parameters like these.

* `Length of movie filename prefix to cut for XML correspondence`**:** Use this field to specify the number of characters to cut off the prefix of the imported movie filename, to match with the XML filename.
* `Length of movie filename suffix to cut for XML correspondence`**:** Use this field to specify the number of characters to cut off the suffix of the imported movie filename, to match with the XML filename.
* `Length of XML filename prefix to cut for movie correspondence`**:** Use this field to specify the number of characters to cut off the prefix of the XML filename, to match with the imported movie filename.
* `Length of XML filename suffix to cut for movie correspondence`**:** Use this field to specify the number of characters to cut off the suffix of the XML filename, to match with the imported movie filename.

## **Outputs**

* Exposures dataset with imported beam shift values

## **Common Next Steps**

* [CTF Estimation](/processing-data/all-job-types-in-cryosparc/ctf-estimation/job-patch-ctf-estimation)
* [Exposure Group Utilities](/processing-data/all-job-types-in-cryosparc/ctf-refinement/job-exposure-group-utilities)


# Motion Correction

## Overview

During imaging in the microscope, a cryo-EM sample is irradiated by the electron beam for typically between one and ten seconds. During this time, the sample does not remain perfectly still.

Drift of the stage, vibration of the microscope, and deformation of the sample ice all contribute to motion of the particles. These effects are all visible as motion blur in the final recorded image. To allow for correction of this effect, microscope images are captured as multi-frame movies (typically approximately fifty frames over the entire exposure). Because each *frame* only captures a short amount of time, the motion blur in each frame is significantly lower than if the entire micrograph was collected at once (i.e., if there was only one frame).

Motion correction is the process by which those frames of a raw movie are aligned and averaged to produce a single-frame micrograph. This process significantly improves signal-to-noise ratio over collecting data in a single frame by reducing the cumulative effect of motion blur.

In addition to causing motion, the beam interacts directly with the sample. The electron beam is a powerful source of radiation, capable of damaging the sample. High-resolution features are especially sensitive to radiation damage. Since the late frames have received the greatest radiation dose, they also tend to have the lowest quality information at high frequencies.

A technique commonly called dose weighting (Grant and Grigorieff, 2015) accounts for the varying information content in each frame by attenuating the high-frequency signal from later frames of movies. All motion-correction jobs in CryoSPARC apply dose weighting. See [the relevant section of this page](#dose-weighting) for more information on the specific forms of dose-weighting applied by each job.

CryoSPARC provides multiple motion correction methods and workflows. In almost all projects, [Patch Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction) should be used in the initial processing steps. It is also used internally by [CryoSPARC Live](/live/about-cryosparc-live) when performing real-time processing. If the final 3D reconstruction is of high quality, [Reference Based Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-reference-based-motion-correction-beta) may provide additional resolution improvement.

## What is motion correction?

Motion correction is the process of algorithmically correcting for motion of the electron microscope stage and the sample ice itself to recover image quality lost by motion blurring.

<figure><img src="/files/XxgcUCqMavM0seJXmnnq" alt=""><figcaption><p>In this movie of intact SARS-CoV-2 virions (EMPIAR 10492), the particles in total move only approximately 10 Å, but loss of contrast due to motion blurring is still clear. Motion correction recovers signal lost due to motion blurring, increasing contrast and improving the quality of the final results. Though not obvious in individual micrographs, the effect of motion correction is most important in improving high-resolution signal quality.</p></figcaption></figure>

Cryo-EM data is collected in the form of movies, which are each a series of individual frames. Since a frame is usually between 0.1 and 0.2 seconds, the detector does not accumulate enough electron dose for clear identification of the target. However, the brief length of time significantly reduces the amount of in-frame motion blur.

<figure><img src="/files/re4IdZJF79teOYrRiA36" alt=""><figcaption></figcaption></figure>

Since the same physical objects create the image in each frame, we can find a shift to apply to each frame (or sub-region of a frame) that results in the greatest agreement among all frames. The ultimate goal of motion correction is to reduce the total blurring in the final particle images that are extracted — what differs between the different methods and implementations is the type of input data they require and the types of motion they are capable of capturing and correcting.

### What causes motion during data collection?

There are two main forces which cause motion during data collection: stage drift and beam-induced ice deformation.

The cause of the former is self explanatory: mechanical effects cause drift of the entire stage and grid during movie collection. This motion is observed as a *shift of the entire image frame*, and is the easiest type of motion to model. Each frame can simply be translated to create the best match between the previous and the next. This is called *Rigid* or [*Full Frame* Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-full-frame-motion-correction), because it produces a movement trajectory of the entire frame as a rigid object. Note that rotation of the grid is generally negligible and is not modeled.

When the sample is irradiated by the electron beam, the thin layer of ice suspended in the grid hole buckles. This three-dimensional movement appears in the movie as anisotropic (i.e., different at different spatial positions) movement of the ice itself. The exact mechanism behind this effect is not fully understood, but it is suspected that the electron beam allows for relaxation of physical stress built up during sample vitrification (Thorne 2020). Modeling this type of motion is more challenging, since in theory small image regions in each frame might move in a different directions. CryoSPARC models anisotropic motion primarily using the [Patch Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction) job, and the theory and method underlying that job is described in the job page.

## What does each Motion Correction job do?

As discussed above, the goal of motion correction is to align each frame such that the particles are in the same position throughout the movie. This way, when they are averaged, the high-resolution information is not lost due to motion blur. For example, consider the movement of particles move throughout this movie:

<figure><img src="/files/pRwMKVbp1dWrUiIHZ8se" alt=""><figcaption></figcaption></figure>

If this movement is not corrected, the particle images would be unacceptably blurry, destroying high-resolution information:

<figure><img src="/files/q9HLJ4QHKnSsoOcB9YjQ" alt=""><figcaption></figcaption></figure>

### Full Frame Motion Correction

The main source of motion for a given particle is typically concerted motion of the entire stage, or frame. It is relatively straightforward to correct for this type of motion by minimizing the difference between one frame and its neighbors. In doing this, each frame is brought into register with the others. This corrects the rigid movement of the entire frame, while leaving the anisotropic movement unmodelled:

<figure><img src="/files/nVnmAdaxVm11y0pJ3qoA" alt=""><figcaption></figcaption></figure>

While this unmodelled movement still results in blurry particles, the results of averaging these aligned frames is already much clearer than if the movie had been collected as a single-frame micrograph:

<figure><img src="/files/xT99VcvSWAmkw03H3ai4" alt=""><figcaption></figcaption></figure>

As this is a relatively simple form of motioncorrection, it was one of the earliest forms introduced. Some examples of early implementations of Full Frame Motion Correction are MotionCor (Li et al. 2013) and Unblur (Grant and Grigorieff 2015).

### Local Motion Correction and Reference Based Motion Correction

Correcting the anisotropic motion of individual particles is more challenging than the full-frame motion for two main reasons. First, and most importantly, single particle images each contain far less signal than an entire micrograph. Second, correcting for anisotropic motion requires knowing the position of particles to begin with. Once good particle location information is available, there are two types of jobs to correct the movement of individual particles in CryoSPARC.

The first is [Local Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-local-motion-correction). In this job, a small patch around each particle is compared in each frame and aligned so as to reduce the total motion, much like the process in Full Frame Motion Correction. The second type is [Reference Based Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-reference-based-motion-correction-beta), which compares a projection of a high-quality 3D Volume to the particle location in each frame to find that particle’s exact location.

In general, with a high quality reference volume, we expect Reference Based Motion Correction to perform better than Local Motion Correction. Reference Based Motion Correction also estimates empirical dose weights (see the Dose weighting section) and is based on Bayesian Polishing (Zivanov et al. 2019). However, the requirement of a high-quality volume is significant. Local Motion Correction does not require a reference, and so can be run early on in the processing pipeline if significant anisotropic motion is observed. Local Motion Correction is based on alignparts\_bfgs (Rubinstein et al. 2015).

### Patch Motion Correction

[Patch Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction) is a fourth motion correction job available in CryoSPARC. It provides a means of correcting anisotropic motion without knowing particle positions. To do this, it models the movement of large patches of the micrograph, then describes that movement using a function called a *spline*. Importantly, this spline can be evaluated at any pixel location to find the trajectory of that pixel during the recording of the movie.

When a particle image must be extracted from a position in a Patch Motion corrected micrograph, the spline function is evaluated at each pixel position to create an aligned average for that pixel. In this way, anisotropic motion is corrected without knowing the particle locations before hand. Therefore, this job typically gives better results than Full Frame Motion Correction, and is the job we recommend for motion correcting any new dataset in CryoSPARC. More information on this algorithm is available in the [Patch Motion Correction job page](/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction).

## Dose weighting

### What is dose weighting?

<figure><img src="/files/V2XYWjshfK7mrXnlZCnX" alt=""><figcaption><p>A comparison of a map made from the first fifteen or last sixteen frames of a set of movies. The same particles are used in the same poses in both images — only the frames used during motioncorrection differ. The data are from EMPIAR 10424 (Nakane et al. 2020). The EER movies were fractionated into 31 frames with an upsampling factor of 2. The same 500 movies were used in each patch motion correction job. The motion correction jobs were run with default parameters, except F-crop set to 1/2. Published poses were used for homogeneous reconstruction of particles extracted from non-dose-weighted micrographs at a box size of 700 px, downsampled to 320 px.</p></figcaption></figure>

During image collection, samples are bombarded with an intense electron beam. This beam damages the fragile bonds in the macromolecule, with the amount of damage increasing as electron dose accumulates. In 3D reconstructions, this is visible as a degradation of the high-resolution signal in later frames of the movie — the high resolution information has been burned away by the electron beam.

To alleviate the worst effects of this radiation damage, the process of dose weighting involves down-weighting the contribution of late frames to the micrographs. In this way, it is possible to both

* retain the low-frequency signal from late frames, which is less damaged by radiation and useful for particle picking
* discard the high-frequency noise from late frames (since there is almost no useful information at these frequencies).

We can plot the dose weights as a series of bar graphs in which the first frame is the topmost bar and the last frame is the lowest bar, and the weight of a frame at a particular resolution is given by the length of the bar.

<div data-full-width="true"><figure><img src="/files/5Kz16FNLeCb9x7QB278B" alt=""><figcaption></figcaption></figure></div>

When plotted this way, the trend aimed for by dose weighting becomes apparent. At low resolution (left panels), all frames are more-or-less equally reliable since the effect of the electron beam is much less noticeable at this resolution. Therefore all frames are treated equally in the final micrograph — they all have a weight of 1.0.

However, radiation rapidly damages information at the highest frequencies (right panels). We therefore want to use only information from the early frames at the highest resolution, so early frames have a weight greater than 1.0 and late frames have a weight less than 1.0.

Visualizing dose weights in this way can become unwieldy when considering all frequencies in a movie. We therefore typically present them as a heatmap instead, where the columns correspond to a resolution, the rows correspond to frames, and the color denotes the weight associated with that frame at that resolution.

<figure><img src="/files/WAyHEp13DXMZqOx7xBud" alt=""><figcaption></figcaption></figure>

### Dose weighting in CryoSPARC

The heatmap above shows an example of the default dose weights applied to movies during motion correction. These default dose weights are calculated in the same way as described by Grant and Grigorieff. These dose weights are based on a model of exponential decay of the signal-to-noise ratio and are applied without any knowledge of the underlying sample or movies.

Default dose weights work well in most cases, but if more information about the system is available a better estimate of radiation damage can be derived. More specifically, if the position of particles in each frame, high-quality pose estimates, and high-resolution reference volumes are available for each particle, it is possible to calculate the correlation between the reference volume and the particle image. From these correlations we can deduce the appropriate dose weights for each frame and resolution.

These calculated dose weights are called **empirical dose weights** and can be calculated by comparing a 3D reference with the particles in the movie. This process is performed, for example, by Bayesian Polishing in RELION (Scheres 2014; Zivanov et al. 2019) and [Reference Based Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-reference-based-motion-correction-beta) in CryoSPARC. More information about the procedure is available in that job page.

## References

1. Thorne, R. Hypothesis for a mechanism of beam-induced motion in cryo-electron microscopy. *IUCrJ* vol. 7 416–421 (2020).
2. Li, X. *et al.* Electron counting and beam-induced motion correction enable near-atomic-resolution single-particle cryo-EM. *Nature Methods* **10**, 584–590 (2013).
3. Grant, T. & Grigorieff, N. Measuring the optimal exposure for single particle cryo-EM using a 2.6 Å reconstruction of rotavirus VP6. *eLife* **4**, e06980 (2015).
4. Zheng, S. Q. *et al.* MotionCor2: anisotropic correction of beam-induced motion for improved cryo-electron microscopy. *Nature Methods* **14**, 331–332 (2017).
5. Nakane, T. *et al.* Single-particle cryo-EM at atomic resolution. *Nature* **587**, 152–156 (2020).
6. Scheres, S. H. Beam-induced motion correction for sub-megadalton cryo-EM particles. *eLife* **3**, e03665 (2014).
7. Zivanov, J., Nakane, T. & Scheres, S. H. W. A Bayesian approach to beam-induced motion correction in cryo-EM single-particle analysis. *IUCrJ* vol. 6 5–17 (2019).
8. Rubinstein JL, Brubaker MA. Alignment of cryo-EM movies of individual particles by optimization of image translations. *J Struct Biol* (2015).

## Motion Correction Jobs

{% content-ref url="/pages/-MR1pPxz8oV8vD7hR-oy" %}
[Job: Patch Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction)
{% endcontent-ref %}

{% content-ref url="/pages/-MR1pK52kx\_tgOuk3GoV" %}
[Job: Full-Frame Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-full-frame-motion-correction)
{% endcontent-ref %}

{% content-ref url="/pages/kNgvuDHqQUT39iBVJjSv" %}
[Job: Local Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-local-motion-correction)
{% endcontent-ref %}

{% content-ref url="/pages/9yzqa8SeDKMwjIEMn7Vt" %}
[Broken mention](broken://pages/9yzqa8SeDKMwjIEMn7Vt)
{% endcontent-ref %}

{% content-ref url="/pages/dtTZBOpS2X7VWo87kp5D" %}
[Job: MotionCor2 (Wrapper) (BETA)](/processing-data/all-job-types-in-cryosparc/motion-correction/job-motioncor2-wrapper-beta)
{% endcontent-ref %}

{% content-ref url="/pages/OdXx0PGmfMOyzBAgHbJN" %}
[Job: Reference Based Motion Correction (BETA)](/processing-data/all-job-types-in-cryosparc/motion-correction/job-reference-based-motion-correction-beta)
{% endcontent-ref %}

## Motion Correction Tutorials

{% content-ref url="/pages/-MNeA4BFZRumxWAGw8dk" %}
[Tutorial: Patch Motion and Patch CTF](/processing-data/tutorials-and-case-studies/tutorial-patch-motion-and-patch-ctf)
{% endcontent-ref %}


# Job: Patch Motion Correction

## At a Glance

Model full-frame and anisotropic motion in movies, and apply dose weighting, to produce motion-corrected micrographs.

## Description

During movie collection, complex 3D deformations of the ice occur. When projected into a flat image, these 3D deformations appear as anisotropic motions of the ice, in which different regions of the image move in different directions. See e.g., (Thorne, 2020) for more details about the mechanism of anisotropic beam-induced motion.

Patch Motion Correction models this anisotropy by tiling a movie into patches and estimating the displacement of each patch in each frame under a motion model that is smooth over both space and time. It then corrects for the estimated motion by shifting each pixel based on its modeled displacement in each frame. Similar approaches are used in software such as MotionCor2 (Zheng et al. 2017) and Warp (Tegunov et al. 2018). Patch motion correction also applies [Dose Weighting](/processing-data/all-job-types-in-cryosparc/motion-correction#dose-weighting).

## Inputs

### Movies

This job requires raw movies in `.mrc` , `.tif` , `.eer`, or `.mrc.bz2` format, typically from an Import Movies job.

## Commonly Adjusted Parameters

### Only process this many movies

This parameter selects `n` movies, randomly to process. It is most useful when working with a subset of movies to assess data quality before committing to a full processing pipeline. If this is left blank, all movies are processed.

To select the same movies every time, set the `Random seed`parameter to a constant value.

### Low-memory mode

Reading movie files from disk is slow. To speed up processing, motion correction jobs process one movie and load the next simultaneously. However, this can lead to GPUs with low memory (i.e., < 16 GB VRAM) to run out of memory. Turning this option on causes the job to wait to load the next movie until it has finished processing the current movie. This slows down the job, but can prevent out-of-memory errors.

### Save results in 16-bit floating point

Save the output micrographs in 16-bit floating point. We recommend that this option is turned on to save space at minimal loss of accuracy. See the [16-bit floating point article](/processing-data/tutorials-and-case-studies/tutorial-float16-support) for more information.

### Output denoiser training data

When this parameter is turned on, Patch Motion Correction will produce the even and odd half-micrographs necessary to train the [Micrograph Denoiser](/processing-data/all-job-types-in-cryosparc/exposure-curation/job-micrograph-denoiser-beta). Note that these half-micrographs do take up a small amount of disk space — the training data pair takes up as much space as the motion-corrected micrograph. For a 4k frame saved in 16-bit format, the default 200 movie training set takes up approximately 12 GB of disk space.

### Num. movies for denoiser training data

This parameter sets the number of movies for which denoiser training data is produced and defaults to 200. Generally, 100 micrographs are sufficient to train the denoiser. By default, more training data is produced in case some micrographs with training data are excluded in, e.g., a [Curate Exposures](/processing-data/all-job-types-in-cryosparc/exposure-curation/interactive-job-manually-curate-exposures) job. If space is a significant concern, or if no micrographs will be excluded before training the denoiser, this parameter can be reduced to the number of micrographs you plan to use when training the denoiser.

### Output F-crop factor:

Reduce the size of the output, aligned micrographs by this factor. If the F-crop (Fourier-crop) factor is 1.0, no cropping is performed. A crop factor of 0.5 downsamples output micrographs by half, discarding the highest-frequency half of the Fourier components, etc. This significantly reduces the file size of the final micrographs, but limits the maximum resolution possible for resulting reconstructions.

Particle images can also be Fourier-cropped (downsampled) at the [Extract from Micrographs](/processing-data/all-job-types-in-cryosparc/extraction/job-extract-from-micrographs) stage. It is often common to initially process downsampled particle images (to speed up particle curation and early reconstructions) and then return to full resolution particle images downstream in final refinements. This type of workflow is made easier if the motion corrected micrographs are not downsampled, and the downsampling is instead only applied at the extraction stage. As such, for a general workflow and non-super-resolution movies, we recommend using the default Output F-crop factor of 1.0 during motion correction. When working with super-resolution movies, we recommend downsampling to the physical detector pixel size (i.e., setting Output F-crop factor to 0.5 for 2x super-resolution movies). Of course, there are some datasets that benefit from the full, super-resolution pixel size, and some which benefit from more aggressive downsampling.

### Start frame and End frame

These parameters allow discarding frames from the start and end of the input movies, before motion correction is applied. This can be helpful in case there are artefacts or issues with early or late frames. For example, some cameras can have issues where they output one or more blank frames at the end of a movie, and these can disturb motion estimation.

### Override knots X

Spline knots control the smoothness of motion trajectories that will be optimized by Patch motion correction (see [Method Details](#method-details) below). CryoSPARC automatically determines the number of knots in each spatial and temporal direction based on characteristics of the input movies (e.g., magnification, dose rate, etc) and it is not generally necessary to change these parameters. These parameters can be used to enforce additional smoothness on motion trajectories (setting knots to a low value) or allow patch motion correction to model additional anisotropy (setting knots to a high value) beyond its automatic settings, if needed.

Importantly, the knot parameters **do not** change the number of patches or size of patches used by Patch Motion Correction. Patches are 500 Å wide and cannot be resized by the user.

## Outputs

### Micrographs

Aligned micrographs for further downstream preprocessing. These micrographs also have rigid motion estimates, which are required for jobs like [Reference-Based Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-reference-based-motion-correction-beta) and [Local Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-local-motion-correction).

### Micrographs incomplete

Micrographs which were not successfully corrected are output in this group. Typically this is due to some error specific to the movie — the Event Log should contain more information on why a particular movie failed.

## Common Next Steps

Typically, micrographs must have their CTF estimated, via e.g., [Patch CTF Estimation](/processing-data/all-job-types-in-cryosparc/ctf-estimation/job-patch-ctf-estimation), before any further processing can occur.

## Method Details

Patch Motion Correction progresses through three main steps for each movie: rigid motion estimation, anisotropic motion estimation, and motion correction.

Rigid motion is estimated in much the same way as [Full-Frame Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-full-frame-motion-correction). More detail is available in that article, but in short, a full-frame trajectory is iteratively optimized to produce the highest degree of correlation between all frames.

<figure><img src="/files/oo7f83dFC5YH8oPQ4sgV" alt=""><figcaption><p>An example micrograph from EMPIAR 10674.</p></figcaption></figure>

Next, the micrograph is split into overlapping 500 Å patches. In order to model anisotropic motion (i.e. deformation of the sample), the motion trajectories of all patches are jointly optimized to produce the highest degree of correlation across all frames of the movie. In this optimization setup, for each patch, each frame (denoted by a Z dimension representing time) is allowed to shift in both X and Y directions.

<figure><img src="/files/pk8SvJPgUEhXQYV3ECR3" alt=""><figcaption><p>An example of a single frame’s patch displacement. An arrow represents the displacement of a patch in a single frame. Note that the length of the arrows are scaled for visibility.</p></figcaption></figure>

The optimization of patch trajectories is constrained using a spline function. The number of knots (i.e. degrees of freedom) in the spline function is automatically selected by CryoSPARC based on the magnification and dose rate of the input movies, and the spline serves to restrict the possible patch trajectories that the optimization can model to ensure that they are smooth over both space and time. Furthermore, after optimization, the spline function allows interpolation of displacement for each pixel, rather than only at the patch level.

<figure><img src="/files/HUjYhyKZHEyKqEFzovFd" alt=""><figcaption><p>X and Y displacement models for a single frame are shown with pink and blue surfaces, respectively. Using this function, each pixel’s displacement can be modeled.</p></figcaption></figure>

As mentioned, the number of spline knots control the number of degrees of freedom used by the motion model in each of the three dimensions (X, Y, and time). A small number of knots causes the model to be smooth, which protects from overfitting but may also prevent the spline functions from fully capturing the dynamics of patch movement. A large number can have the opposite effect. CryoSPARC automatically selects the number of knots.

The full motion model, optimized using data across all frames, provides an estimate of the anisotropic motion during the entire movie:

<figure><img src="/files/U6eG6SxBpB3FfFO486vF" alt=""><figcaption></figcaption></figure>

which in turn can be used to model each pixel’s trajectory through time and space:

<figure><img src="/files/6WNpWWtdHc1bC8L0PA6S" alt=""><figcaption><p>Each pixel’s displacement varies as it moves through time, and so do the spline functions.</p></figcaption></figure>

After motion estimation, the final spline function provides the optimal prediction of the shift of any given pixel for each frame. This prediction is used to shift each pixel in each frame by the modeled displacement to produce a single, motion corrected averaged micrograph from which particles can be picked and extracted. Note that this is distinct from other motion correction jobs, like Local Motion Correction, that model and correct the motion of individual particles and produce extracted particles rather than a complete micrograph.

## Example Workflow

{% content-ref url="/pages/-MNeA4BFZRumxWAGw8dk" %}
[Tutorial: Patch Motion and Patch CTF](/processing-data/tutorials-and-case-studies/tutorial-patch-motion-and-patch-ctf)
{% endcontent-ref %}

## References

1. Robert Thorne, “Hypothesis for a Mechanism of Beam-Induced Motion in Cryo-Electron Microscopy,” *IUCrJ*, 2020, <https://doi.org/10.1107/S2052252520002560>.
2. Zheng, S. Q. *et al.* MotionCor2: anisotropic correction of beam-induced motion for improved cryo-electron microscopy. *Nature Methods* **14**, 331–332 (2017).
3. Tegunov, D. & Cramer, P. Real-time cryo-EM data pre-processing with Warp. *bioRxiv* (2018) doi:[10.1038/s41592-019-0580-y](https://doi.org/10.1038/s41592-019-0580-y).


# Job: Full-Frame Motion Correction

## At a Glance

<figure><img src="/files/5QxTolQPTllbjvNIsn0m" alt=""><figcaption></figcaption></figure>

Perform rigid motion correction of movies.

## Description

Full-frame motion correction models and corrects movement of each full frame of the movie. It therefore cannot correct anisotropic movement, such as movement due to ice doming during movie collection. Full-frame motion correction also applies [Dose Weighting](/processing-data/all-job-types-in-cryosparc/motion-correction#dose-weighting).

See [Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction) for more details about the different types of motion correction available in CryoSPARC.

## Inputs

### Movies

Full-Frame Motion Correction requires raw movies in `.mrc` , `.tif` , `.eer`, or `.mrc.bz2` format, typically from an [Import Movies](/processing-data/all-job-types-in-cryosparc/import/job-import-movies) job.

## Commonly Adjusted Parameters

### Only process this many movies

This parameter selects the first `n` movies to process, rather than the entire input stack. It is most useful when working with a subset of movies to assess data quality before committing to a full processing pipeline.

If collection parameters changed during the collection, it may be better to use an [Exposure Sets Tool](/processing-data/all-job-types-in-cryosparc/utilities/job-exposure-sets-tool) job to select a *random* subset instead.

### Low GPU memory

Reading movie files from disk is slow. To speed up processing, motion correction jobs process one movie and load the next simultaneously. However, this can lead to GPUs with low memory (i.e., < 16 GB VRAM) to run out of memory. Turning this option on causes the job to wait to load the next movie until it has finished processing the current movie. This slows down the job, but can prevent out-of-memory errors.

### Save results in 16-bit floating point

Saves the output micrographs in 16-bit floating point. We recommend that this option is turned on to save space at minimal loss of accuracy. See the [16-bit floating point article](/processing-data/tutorials-and-case-studies/tutorial-float16-support) for more information.

### Start frame and End frame

These parameters allow discarding frames from the start and end of the input movies, before motion correction is applied. This can be helpful in case there are artefacts or issues with early or late frames. For example, some cameras can have issues where they output one or more blank frames at the end of a movie, and these can disturb motion estimation.

## Outputs

### Micrographs

Aligned micrographs for further downstream preprocessing. These micrographs also have rigid motion estimates, which are required for jobs like Reference-Based Motion Correction and Local Motion Correction.

### Micrographs incomplete

If alignment of any micrographs failed, the failed micrographs are collected in this output. If the micrographs failed because the GPU ran out of memory during processing, re-processing these movies after freeing some GPU memory may successfully resolve the issue. Otherwise, these movies may be corrupted.

## Common Next Steps

Typically, micrographs have their CTF estimated by Patch CTF Estimation before any further processing occurs.

## Recommended Alternatives

In almost all cases, [Patch Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction) will outperform Full-Frame Motion Correction. Patch Motion Correction accounts for both rigid motion (as with Full-Frame) and the local, anisotropic motion which occurs in all cryoEM movies. We recommend that users prefer Patch Motion Correction unless they have very low SNR movies or very sparse particle distribution.


# Job: Local Motion Correction

Local motion correction.

## At a Glance

Track and correct the motion of individual particles in movies.

## Description

Local Motion Correction job corrects beam-induced motion on a per-particle basis, taking full-frame-motion corrected movies and particle pick locations as inputs. A square patch around each particle is extracted from each frame of the movie. The trajectory of these patches is modeled and then used to extract a motion-corrected patch for each particle and each frame. Local motion correction also applies [Dose Weighting](/processing-data/all-job-types-in-cryosparc/motion-correction#dose-weighting). Local motion correction is based on alignparts\_lbfgs (Rubinstein et al. 2015).

Two versions of Local Motion Correction are available - Local Motion Correction runs on a single GPU, while Local Motion Correction (Multi-GPU) uses multiple available GPUs to parallelize processing.

This job does *not* require a high-quality 3D reference volume (unlike Reference Based Motion Correction), but it does require well-centered particle picks. A typical workflow would be to first use patch motion correction on the raw movies, perform picking, and then re-perform motion correction with Local motion correction.

See [Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction) for more details about the different types of motion correction available in CryoSPARC.

## Inputs

### Movies

Local motion correction requires raw movies in `.mrc` , `.tif` , `.eer`, or `.mrc.bz2` format, typically from an Import Movies job.

### Particles

Particles should be well-centered and relatively clean of junk or empty picks, most often from a 2D Classification or any 3D refinement job.

## Commonly Adjusted Parameters

### Only process this many movies

This parameter selects the first `n` movies to process, rather than the entire input stack. It is most useful when working with a subset of movies to assess data quality before committing to a full processing pipeline.

If collection parameters changed during the collection, it may be better to use Exposure Sets to select a *random* subset instead.

### Save results in 16-bit floating point

Saves the output micrographs in 16-bit floating point. We recommend that this option is turned on to save space at minimal loss of accuracy. See the [16-bit floating point article](/processing-data/tutorials-and-case-studies/tutorial-float16-support) for more information.

### Override e/A^2/frame

When `Apply exposure weighting` is enabled, this parameter can be used to override the exposure level of the movie data, in case the level was not correctly set at import time. In contrast to the Import Movies job's `Total exposure dose (e/A^2)` parameter, Local Motion Correction's parameter must be specified *per-frame*. Exposure weighting is based on the reference curves from Grant and Grigorieff.

### Extraction box size

The side length of the extraction box, in pixels. The general rule of thumb is twice the width of the particle, but more precise guidance is given by Rosenthal and Henderson as

$$
\mathrm{Box\ Size\ (\AA{})} = d\ +\ (2\lambda{}\times{}\frac{\delta{}}{r})
$$

where

* *d* is the particle diameter in Å,
* λ is the electron wavelength (approximately 0.020 Å and 0.025 Å for 300 kV and 200 kV microscopes, respectively),
* δ is the defocus (in **Å**, not μm), and
* *r* is the expected resolution of the final reconstruction (the average CTF fit would be a good starting value if data quality is not known)

The first term of this equation is simply the particle diameter, an obvious lower limit on box size, since the entire particle must fit in the box. The second term models the displacement of signal by the CTF. With higher defocus, higher frequency information is displaced by a greater distance, hence the direct dependence on defocus and inverse dependence on final resolution, since higher resolution decreases *r*.

### Fourier crop to box size (pix)

Extracted, motion-corrected particle images will be Fourier cropped to this box size.

## Outputs

### Particles extracted

Since particle locations are already known and motion of the entire movie is not modeled, this job outputs motion corrected particle images rather than whole micrographs.

### Diagnostic plots

Local Motion Correction also produces a plot of the modeled movement of each particle:

<figure><img src="/files/4FNweiKLrlTpJsyXCi6B" alt=""><figcaption></figcaption></figure>

The local motion plot shows the motion modeled for each particle patch. The trajectories are scaled by 40x to be visible on the micrograph scale.

The grey line shows the raw trajectory. Because single-particle data typically has a very poor signal-to-noise ratio, these trajectories are very noisy and likely not a good representation of the true movement of the particle. The per-frame trajectory is therefore smoothed out to produce the final result. This smooth trajectory is in red.

## Common Next Steps

Particle images from Local Motion Correction are suitable for use in a variety of jobs, including 3D Refinements and heterogeneity analysis techniques.

## Recommended Alternatives

If a high-quality reconstruction is available for the dataset, [Reference Based Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-reference-based-motion-correction-beta) will likely perform better than Local Motion Correction. Like Local Motion Correction, it directly models particle trajectories. However, Reference Based Motion Correction also calculates empirical dose weights (rather than use the reference curves), which can lead to improved final reconstruction quality.

## References

1. Rubinstein JL, Brubaker MA. Alignment of cryo-EM movies of individual particles by optimization of image translations. *J Struct Biol* (2015).
2. Rosenthal, P. B. & Henderson, R. Optimal Determination of Particle Orientation, Absolute Hand, and Contrast Loss in Single-particle Electron Cryomicroscopy. *Journal of Molecular Biology* **333**, 721–745 (2003).
3. A table of electron wavelengths for a given accelerating voltage can be found here: <https://www.jeol.com/words/emterms/20121023.071258.php#gsc.tab=0>
4. Grant, T. & Grigorieff, N. Measuring the optimal exposure for single particle cryo-EM using a 2.6 Å reconstruction of rotavirus VP6. *eLife* **4**, e06980 (2015).


# Job: MotionCor2 (Wrapper) (BETA)

How to use the wrapper for MotionCor2 available in CryoSPARC.

## At a Glance

Perform motion correction using MotionCor2 as a wrapper within CryoSPARC.

## Description

MotionCor2 models the anisotropic motion of electron microscopy movies using a smooth deformation model. More details are available in Zheng, 2017.

Please ensure that you review the MotionCor2 License Terms before using it in your project.

### MotionCor2 License Terms

> CryoSPARC does not distribute MotionCor2 binaries. Please ensure you have your own copy of MotionCor2 installed under the terms of the MotionCor2 Non-Commercial Software License Agreement available at: <https://docs.google.com/forms/d/e/1FAIpQLSfAQm5MA81qTx90W9JL6ClzSrM77tytsvyyHh1ZZWrFByhmfQ/formResponse>.\
> \
> For-profit users must contact David Agard for licensing information prior to download. Structura Biotechnology Inc. makes no warranty regarding MotionCor2.

## Inputs

### Movies

MotionCor2 requires movies as an input, typically from an Import Movies job.

## Commonly Adjusted Parameters

### Patch Value

MotionCor2 requires users to input the number of patches in the X and Y axes. Note that this setting is distinct from Patch Motion Correction’s `knots` settings.

## Outputs

### Micrographs

MotionCor2 outputs motion-corrected micrographs.

## Common Next Steps

Micrographs typically require CTF estimation in [Patch CTF Estimation](/processing-data/all-job-types-in-cryosparc/ctf-estimation/job-patch-ctf-estimation) for further downstream processing.

## References

1. Zheng, S. Q. *et al.* MotionCor2: anisotropic correction of beam-induced motion for improved cryo-electron microscopy. *Nature Methods* **14**, 331–332 (2017).


# Job: Reference Based Motion Correction (BETA)

## At a Glance

Use a high-quality reference volume, particle poses, and particle positions to estimate per-particle movement trajectories and empirical dose weights.

## Description

Reference-based motion correction is an extension of [Patch Motion Correction](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction). Using known particle poses and positions, precise movement trajectories can be calculated for each particle. In addition, the effect of radiation damage during an exposure can be empirically measured and accounted for by weighting. In some cases, these procedures yield a significant improvement in final map quality.

The concepts and method in CryoSPARC’s Reference Based Motion Correction are inspired by **Bayesian Polishing** (Zivanov, Nakane & Scheres, 2019). CryoSPARC’s implementation includes a new method for hyperparameter optimization, is multi-GPU accelerated and optimized, and includes support for multiple reference volumes, thereby enabling simultaneous motion correction for particles from different conformations/species. In addition, patch motion correction can take as input the empirical dose weights estimated by Reference Based Motion Correction on a different dataset.

This job type has an accompanying tutorial video:

{% embed url="<https://www.youtube.com/watch?v=gnM_IvJShwY>" fullWidth="true" %}

## Inputs

{% hint style="info" %}
In CryoSPARC v4.4, the Exposures input was called Movies. For the purposes of this job, the two names are interchangeable.
{% endhint %}

{% hint style="info" %}
To properly match a particle with its given reference, Reference Based Motion Correction accepts sets of Particles, Volumes, and Masks. These sets are given in separate numbered inputs. To include more than one set, increase the `Number of Reference Volumes` parameter.
{% endhint %}

<div data-full-width="true"><figure><img src="/files/gmJZEisXqGZAjHTWSbet" alt=""><figcaption></figcaption></figure></div>

### Exposures

The connected exposures must have rigid motion estimates and background estimates. A [Patch Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction) job provides these estimates.

### Volumes

The connected reference volumes must have half-maps and a mask. Jobs such as [Homogeneous Refinement](/processing-data/all-job-types-in-cryosparc/3d-refinement/job-homogeneous-refinement) include a mask with the volume output, but the user can always provide a mask of their own by connecting an input to the optional “Static mask N” slots.

{% hint style="info" %}
Reference-based motion correction supports heterogeneous datasets. A parameter titled `Number of reference volumes` (default 1) can be increased to allow multiple particle stacks and reference volumes to be connected.
{% endhint %}

### Particles

It is recommended that you provide particles from the same refinement as the reference volume, and that the refinement job had the `minimize over per-particle scale` switch on.

### Hyperparameters

If you wish, you can connect dose weights and/or motion hyperparameters from a previous job into the `hyperparameters` input group. If you do so, the job will use the supplied motion hyperparameters and/or dose weights instead of recomputing them. If re-processing the same dataset, or processing a separate dataset with similar collection circumstances, the hyperparameters will likely be transferrable.

## Commonly Adjusted Parameters

### Final processing stage

This parameter can be used to stop the job early. For example, after computing motion hyperparameters or dose weights.

### Save results in 16-bit floating point

Turning this setting on will cause the motion-corrected particles to be written to disk in half precision (float16 format, see the [guide page](https://guide.cryosparc.com/processing-data/tutorials-and-case-studies/tutorial-float16-support/~/overview) for more information). Though off by default, this is not known to harm subsequent refinement quality in most cases, and reduces the disk space consumed by 50%. Its use is encouraged.

### Override: EER number of fractions

Normally, the reference-based motion correction job will divide an EER file into a number of fractions that was specified when the movies were imported. This parameter allows the fraction count to be overridden. This parameter can only be used if no frames were discarded in patch motion correction (through the use of the start and end frame parameters).

### Recenter particles

If this parameter is active, then the input particles will be re-centered (their pick locations on the movie will be adjusted) based on their optimized shifts from the upstream refinement.

### Skip movies with wrong frame count

If some input movies have a different number of frames from the rest, the job will fail. If the `skip movies with wrong frame count` switch is on, then the most common frame count will be assumed to be correct and all movies that don't have that frame count will be discarded by the job.

### Hyperparameter search thoroughness

The number of rays that are searched is controlled by the `hyperparameter search thoroughness` parameter, which has 3 options: Fast, Balanced, and Extensive. The fast setting is usually sufficient, and completes in the shortest amount of time. The other settings use more rays, at the expense of more computation time. For an explanation of what these rays represent, see the [Hyperparameter Search section](#hyperparameter-search).

### Maximum total prior strength

The parameter `maximum total prior strength` limits the total strength of the priors. To determine the necessary total prior strength, monitoring trajectory activity and the cross-validation score is helpful. **Trajectory activity** is the average (across the micrographs used in the hyperparameter search) of the per-micrograph 75th percentiles of trajectory length (relative to rigid motion). If, on the last iteration, the trajectory activity hasn’t reached a value very close to zero, the maximum total prior strength may need to be increased.

<div data-full-width="true"><figure><img src="/files/eo7IBMBkqFnoi2p3gNdO" alt=""><figcaption></figcaption></figure></div>

### Fraction of FCs to use for alignment

This parameter determines how many of the Fourier components are used when computing the trajectories, with the remainder being used for cross-validation (see [the theoretical overview section](#theoretical-overview) for details). The default setting usually does not need to be changed.

### Target number of particles

Only a subset of the overall dataset is needed to estimate hyperparameters. The `Target number of particles` parameter sets the number of particles to be used. Micrographs are randomly selected from the dataset one-by-one until they have at least this many particles, or the entire dataset is used.

### Overriding hyperparamter optimization

If you wish to skip the hyperparameter optimization stage entirely, you can do so either by connecting hyperparameters from a previous job, or by manually entering numerical values in the three `override:` parameters.

You must supply either all three of these overrides, or none of them.

### Use all Fourier components

Hyperparameter search only uses the lower-frequency Fourier components when computing trajectories. By default, the final iteration uses all frequency components instead. This improves the quality of the final step and is usually best, but can be turned off by turning `Use all Fourier components` off.

### Fourier-crop to box size

The `Fourier-crop to box size` parameter can be used to reduce the pixel resolution of the output particles by Fourier-space cropping. By default, the particles are extracted using the raw pixel size of the movies (including the upsampling factor, in the case of EER movies) and whatever box size is necessary for the extracted particles to have the same physical extent as the reference volume.

{% hint style="info" %}
The default box size and resolution do not necessarily equate with the motion corrected pixel size used in earlier processing steps (e.g., super-resolution movies).
{% endhint %}

### Number of GPUs

Increasing the `Number of GPUs` parameter can speed up processing. Good performance scaling to more than 3 GPUs usually requires a reasonably modern and fast CPU (e.g., 3rd generation Intel Xeon scalable, AMD Epyc Rome, etc).

Although motion correction calculations are performed on GPUs, a fast CPU is necessary to load data into the GPU. Thus, a given configuration of GPUs may be “too fast” for a given CPU, which would result in GPUs being occupied but not performing at their best.

### GPU oversubscription memory threshold

Any GPUs with more VRAM than the `GPU oversubscription memory threshold` will work on two micrographs at a time instead of one. This can speed up processing, but increases the demand on the CPU. Setting this greater than or equal to GPU VRAM will force a single movie per GPU.

### In-memory cache size

The `in-memory cache size` parameter controls how much RAM is set aside for caching data in the hyperparameter estimation step. This parameter should be set between 60% and 80% of your machine’s RAM, preferably lower unless the machine has more than 256 GB of RAM.

### Slicing GPU also computes trajectories

Normally, the fastest available GPU serves two simultaneous roles: it is responsible for creating particle references by projecting the reference volume through Fourier-space slicing, and it also acts as one of the workers computing trajectory estimates. In problems that have very high VRAM requirements, this can cause the job to fail due to insufficient GPU memory. Turning this switch off will isolate the first GPU for computing references only, thereby reducing VRAM pressure on that GPU. However, doing so also means that the job cannot run unless it is assigned at least two GPUs.

## Outputs

### Empirical Dose Weights

Most cryo-EM motion correction methods, including [Patch Motion Correction](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction), use a dose-weighting scheme predicted from the physics of beam-induced radiation damage along with experimental data on a well-characterized specimen (Grant & Grigorieff, 2015). By contrast, Reference-Based Motion Correction calculates empirical dose weights, on a per-dataset basis, based on the Fourier Cylinder Correlation, or “FCC” (Zivanov, Nakane & Scheres, 2019). The FCC is a measure of how well the aligned frames correlate with the reference volume projections as a function of frame number and spatial frequency. First, the reference volume is projected using the particle’s pose. Then, for each frame, the correlation between the reference projection and the frame image is calculated at each resolution.

During reconstruction, particle images contribute information to the volume across all frequencies, and the images themselves are averages of the patches from movie frames. Empirical dose weighting allows for these sums to be weighted by the “quality” of the particle, as measured by each frame’s correlation with the reference volume at each frequency. These dose-weights are calculated by fitting a model to the FCC, then normalizing each column (i.e., each resolution).

<figure><img src="/files/zB0cZrU2p0TlRHqmUgje" alt=""><figcaption></figcaption></figure>

The first frame has the least radiation damage, and so, for a perfectly static sample, it is theoretically the best source of high-resolution information. However, it is somewhat common for the first frame to exhibit poor correlation at high frequencies due to initial beam-induced motion. In these cases it’s best to trust a slightly later frame (e.g. 2 or 3) for the most high frequency detail. Empirical dose weights account for this; we have found that this effect is responsible for a significant proportion of the typical resolution improvement from the reference motion job overall.

#### Using empirical dose weights in patch motion correction

Also as of CryoSPARC v4.4, Patch Motion Correction has a new optional input for dose weights.

<figure><img src="/files/nPkY4bL32Xiv166Gv3Y0" alt="" width="416"><figcaption></figcaption></figure>

A hyperparameter output group from a reference-based motion correction job can be connected here to use the computed empirical dose weights instead of the standard dose-weighting curve. This might be of use if, for example, there are several datasets to process which were collected at the same time under the same conditions. Since the empirical dose weight computation is sometimes a significant portion of the overall benefit of doing reference-based motion correction, and since reference-based motion correction is quite computationally expensive, it may be possible and convenient to capture some of the benefit at much lower cost in this fashion.

### Motion corrected particles

The final stage of processing shows an overall progress bar and prints out a few example diagnostic plots. The following pair of plots is generated for the first 20 movies processed. After 20 movies, processing continues, but no further plots are generated (refer to the progress bar at the top of the log checkpoint to see overall progress).

<figure><img src="/files/9ITAvolekt8dTrRRILdt" alt=""><figcaption><p>A schematic overview of a micrograph showing particle locations and trajectories (axis labels are in pixels)</p></figcaption></figure>

<figure><img src="/files/Z2Wb59J0cAZwmz0v29rS" alt=""><figcaption><p>Example motion-corrected particles.</p></figcaption></figure>

## Common Next Steps

Particle images from Reference Based Motion Correction are typically used toward the end of analysis in final refinements, such as Non-Uniform or Local Refinements.

## Theoretical overview

Following (Zivanov, Nakane & Scheres, 2019), Reference Based Motion Correction proceeds as follows:

1. For each particle in the input dataset, a synthetic reference image is created by projecting the reference volume in the particle’s pose and applying simulated CTF corruption to the resulting 2D reference.
2. A patch is extracted around the particle’s pick location from each frame in the movie it came from.
3. Each of these patches is then assigned a shift; the set of shifts across all frames makes up the particle trajectory.
4. The optimal trajectory is computed by finding the set of shifts which minimizes the error between the reference image and the shifted patches.
5. Particle images are reconstructed from frames by applying the optimal trajectories and dose weights.

<figure><img src="/files/wuGV0UneBCWw5fgjnKhD" alt=""><figcaption><p>An example of realistic trajectories. Each black dot is a particle pick location, and the tail associated with it is the motion estimate for that particle over the length of the exposure.</p></figcaption></figure>

### Regularization

Due to the particularly low signal to noise ratio present in individual movie frames, the procedure just described would naturally overfit to noise - causing wild and nonsensical trajectories. To mitigate this and following (Zivanov, Nakane & Scheres, 2019), the reference-based motion correction job uses two kinds of regularization: the spatial and acceleration priors.

The spatial prior penalizes candidate trajectories that exhibit low spatial correlation; in other words, the spatial prior encourages solutions where the trajectories of nearby particles are similar to each other. This prior has two parameters that tune its behaviour: an overall strength parameter (how strongly to penalize non-spatially-correlated trajectories) and a correlation distance (over what distance do we expect the trajectories to be similar).

The acceleration prior penalizes trajectories that have high acceleration (i.e. non-smooth trajectories). This prior has one tuning parameter: the overall strength (how strong of a penalty to apply to non-smooth trajectories).

Together the three prior parameters are called the **hyperparameters** of this motion estimation method. If the priors are too strong, the output will have the trajectories set to zero at every frame because the method is ignoring the data and producing a set of trajectories that satisfy the priors. If the priors are too weak they will not achieve their goal, and the method will simply align the references to the noise in the data rather than the signal.

If hyperparameters are not supplied to the job, it will estimate them using the following method.

1. For each particle, two references are created: one from the half-map that the particle contributed to, and one from the opposite half map. This allows for downstream cross-validation, similar to the analysis performed for non-uniform refinement.
2. For a given set of hyperparameters, trajectories are estimated using the low-frequency data from the particle’s half-map.
3. To assess the quality of a set of hyperparameters, the particle trajectories are applied only to the high-frequency part of the data. The resulting corrected images are compared to the opposite half-map. Since high-frequency noise from a particle in half-set A should not correlate with high-frequency signal in half-map B, we can trust the high-frequency correlations from this comparison. This measurement is the cross-validation score, and a more negative number indicates better agreement between the two half-sets.
4. The set of hyperparameters that yields the lowest total cross-validation score is deemed the best. Said another way, the trajectory which correlates best with the ***opposite half map*** is considered best.

#### Hyperparameter search

The following hyperparameter search method is designed to avoid poor selections on a wide range of test datasets.

<div data-full-width="true"><figure><img src="/files/KDok8QFPN2R87HERj5nY" alt=""><figcaption></figcaption></figure></div>

The hyperparameter search is done in a cylindrical coordinate system of the 3-dimensional hyperparameter space. The hyperparameters consist of the total prior strength **r**, acceleration/spatial prior balance **Θ**, and spatial correlation distance **z.** At the start of the search, a number of “rays” (each of which have a fixed z and theta) are created at predetermined positions. During each iteration, the search proceeds outwards along each ray. Once the hyperparameters determined by a ray result in no particle motion, that ray is retired since further increasing the prior strength will have no effect.

## Sample results

### EMPIAR-10061 Beta-galactosidase

In our testing an improvement in FSC resolution of about 0.2 Å is common.

In the following images, reference motion correction is compared against [Patch Motion Correction](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction) on EMPIAR-10061 (beta-galactosidase). In the below images, the blue mesh is Patch Motion, the red is Reference Based Motion Correction. The model (PDB 6DRV) is for illustrative purposes only and has not been refined against the improved map.

<figure><img src="/files/BHqusCKsOhhk503C6Nnr" alt=""><figcaption></figcaption></figure>

In the following images of the same maps, gray is the result from Patch Motion, while cyan is the result from Reference Based Motion Correction.

<figure><img src="/files/h39FHuoHRubuVYvCD0yR" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/aEvXCeqjUs2KYC0ZIcKy" alt=""><figcaption></figcaption></figure>

Finally, for this dataset, we can see from overlayed FSC curves that within reference-based motion correction, the trajectory optimization and empirical dose weighting both contribute to the improvement in resolution, with dose weighting providing a slightly larger part of the improvement. This finding means that for subsequent similar dataset collections, a sizeable improvement could be had by reusing the dose weights estimated from this data in a new patch motion correction job.

<figure><img src="/files/5eGPnSuYrsgRxtKmbCXi" alt=""><figcaption></figcaption></figure>

### EMPIAR-10261 Nav1.7 Ion Channel

The following plots compare the FSC curves from [Patch Motion Correction](https://guide.cryosparc.com/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction) versus reference-based motion correction on the heterogeneous EMPIAR-10261 (Nav1.7 ion channel) dataset. In this case, two volumes were connected as input as both the open and closed conformations of the channel are present in the data. The results show that resolutions for both classes improved.

<div><figure><img src="/files/dJW1lw27mHhcFaNdPAgF" alt=""><figcaption><p>Class 1, patch motion correction</p></figcaption></figure> <figure><img src="/files/2LU8XBAbrJmhZ5RkPmUZ" alt=""><figcaption><p>Class 2, patch motion correction</p></figcaption></figure></div>

<div><figure><img src="/files/Fh72VMwJdpgWLZ4ML84K" alt=""><figcaption><p>Class 1, reference-based motion correction</p></figcaption></figure> <figure><img src="/files/ya4LAAr3kxHPKe2axLzu" alt=""><figcaption><p>Class 2, reference-based motion correction</p></figcaption></figure></div>

## References

1. Jasenko Zivanov, Takanori Nakane and Sjors H. W. Scheres (2019). A Bayesian approach to beam-induced motion correction in cryo-EM single-particle analysis. IUCrJ, 6, 5-17.
2. Timothy Grant and Nikolaus Grigorieff (2015). Measuring the optimal exposure for single particle cryo-EM using a 2.6 Å reconstruction of rotavirus VP6. eLife 4:e06980.


# CTF Estimation

[*Jump to the CTF Job Types*](#ctf-estimation-jobs)

## At a Glance

The Contrast Transfer Function (CTF) models the effect of defocus and microscope aberrations on single particle images. These effects must be corrected before the images can be used to reconstruct a 3D Volume.

This article contains a broad overview of what the contrast transfer function is and where it comes from. Advanced pages provide more detail on waves, contrast, and aliasing. These topics are important both for a deeper theoretical understanding of the technique and also for practical considerations, such as the choice of box size during extraction. Finally, there is a list of useful external resources for interested readers at the end of this page.

{% hint style="info" %}
Developing even an incomplete understanding of the sources of contrast in cryo-EM is a significant undertaking — we intend this section to serve as a reference over the course of one’s cryo-EM education rather than a prerequisite to performing one’s first CTF job.
{% endhint %}

## The Contrast Transfer Function

### The CTF and the PSF

When modeling image aberrations, it may be most intuitive to consider the question, “what does the image of a single point look like when imaged by this system”. This concept is known as the Point Spread Function (often abbreviated PSF). In theory, one might expect that an image of a single point would be a simple projection of that point into a 2D image:

<figure><img src="/files/M0nmmAsopzMqQuRalam9" alt="On the left, the &#x22;object&#x22; is a blue dot. On the right, the &#x22;image&#x22; is a red dot. An arrow labeled &#x22;Point Spread Function&#x22; points between the two." width="563"><figcaption></figcaption></figure>

However, various imperfections and complications in electron microscopes means this is not the case. The deviations of the true image from this “ideal” projected point are collectively called *aberrations*. The largest aberration is due to the fact that images are collected out of focus (also known as “with defocus”) to improve contrast (see [Contrast in Cryo-EM](/cryo-em-foundations/image-formation/contrast-in-cryo-em)). The sum total of these aberrations result, generally for cryo-EM, in a point spread function that consists of oscillating bands of dark and light rings around a single point:

<figure><img src="/files/Pq7PtKdwO9ePlQrPC6YI" alt="The red dot at the right now has oscillating rings of red and white around it." width="563"><figcaption></figcaption></figure>

To model what would happen for an arbitrary object shape, we could break apart the object into a set of many individual points, and apply the PSF at each point to determine the resulting image.

<figure><img src="/files/12p3yYqNR9QwgspvUFx1" alt="At the top, a continuous object on the left has an image with rings around it on the right. At the bottom, the object is split into points, each of which has rings around them in the image." width="563"><figcaption></figcaption></figure>

If the object were split into infinitely many points, applying the point spread function to each point would produce an identical image. This operation is called a convolution, and can be computationally expensive. However, a mathematically useful property of convolution is that it is equivalent to a simple multiplication when working in Fourier space. Thus, to simplify convolution, we can:

1. take the Fourier transform of our object,
2. multiply it by the Fourier transform of the point spread function,
3. and then perform an inverse Fourier transform on the result.

This process produces the image we would expect to see in the microscope.

<figure><img src="/files/nEz05tbip4SeniDKd4aV" alt="This image is a flowchart. At the left, we see the structura logo (labeled &#x22;object&#x22;) and the PSF. Both are Fourier transformed. These Fourier transforms are multiplied together, then the inverse Fourier transform is taken. This results in the final image: the Structura logo with rings around each part."><figcaption></figcaption></figure>

The Fourier transform of the point spread function is so useful and commonly used that it has its own name: the Contrast Transfer Function (CTF). In addition to its usefulness in convolving the object and the point spread function, the contrast transfer function is also much easier to estimate directly from image data than the point spread function.

## The CTF in Practice

The CTF has several practical implications on the processing of cryo-EM data.

#### Particle visibility

<figure><img src="/files/VqF5E0lllpgqtlMFzLvS" alt="A grid of images. The left images have zero defocus, the right images have -2 micron defocus. The top row shows a plot of the CTF at these defoci. The middle images show simulated particles with these defoci. The bottom images show simulated particles plus noise. The images with greater defocus are much easier to see, especially with the added noise."><figcaption></figcaption></figure>

Images collected at or near focus will have very little contrast at low resolutions because protein and buffer scatter electrons with approximately the same intensity (see [Contrast in Cryo-EM](/cryo-em-foundations/image-formation/contrast-in-cryo-em) for more details on this topic). This makes them difficult to pick against a noisy background.

For example, consider the simulated images above. The image collected with no defocus is faintly visible without noise, but becomes almost impossible to see with even modest amounts of noise. Compare this to the image simulated with significant (2 µm) defocus. This level of defocus introduces contrast at low frequencies, making the particle clearly visible even when noise is present. However, the image is more obviously corrupted by oscillations in contrast at high frequencies (visible in the simulated image as black and white rings).

Most, but not all, of the image corruption can be modeled and recovered computationally. Thus, micrographs are typically collected with the least defocus for which particle images can still be reliably picked against the noisy background.

#### Zero crossings

<figure><img src="/files/cKWBlp8kviXw1r8itYdX" alt="A plot of the CTF is shown. Three frequencies are highlighted with points. Below the plot, examples of a wave with this frequency and the result of applying the CTF to this wave are shown."><figcaption></figcaption></figure>

As the contrast transfer function oscillates between -1 and 1, it crosses 0 several times. Frequencies for which the contrast transfer function is 0 have no contrast — they are absolutely invisible in the image.

If all images were collected at the same defocus, all of the images would have 0 contrast at the same spatial frequencies. This would result in a specific band of frequencies having little to no contrast, making the map useless. Data is therefore collected over a range of defocus values such that, taken together, each frequency is represented in a sufficient number of particle images to be properly accounted for.

#### Signal delocalization

<figure><img src="/files/02mdoTbcoISwydvkZxyr" alt="The same particle image is simulated with 0 and -1.5 micron defocus. A zoomed region of the particle shows that the defocused particle has signal in a region that is empty in the no-defocus image."><figcaption></figcaption></figure>

Finally, collecting images with defocus delocalizes signal away from its true position. In the simulated comparison above, note that the high-frequency features (loops in the fabs, helical pitch, etc.) are clearly visible in the left images, where no CTF is applied. Note, of course, that images like this are impossible to collect, since with no defocus, they would not have any contrast. On the right, these same high-frequency features have spread away from their true positions and now overlap, making the image impossible to directly interpret.

This effect is more significant at higher frequencies and at higher defocus values. Particle images which were collected at high defocus and which refine to high resolutions must therefore be extracted with a larger box size to capture information moved away from the particle center.

### Sample Flatness

Cryo-EM samples are never perfectly "flat". Particles tend to concentrate near the air-water interfaces prior to the sample being frozen, and the ice surface itself is often nonplanar. Recalling that defocus affects the contrast transfer function, this means that a single image can contain particles with different defoci and therefore different contrast transfer functions. CryoSPARC provides a patch-based contrast transfer function estimation method in the [Patch CTF Estimation](/processing-data/all-job-types-in-cryosparc/ctf-estimation/job-patch-ctf-estimation) job type which examines many different areas in the micrograph to compute a "defocus landscape", to combat this issue. Patch CTF requires no prior information about particle locations within the micrograph, and can be used immediately after motion correction. It can even work on tilted samples without knowing about the tilt beforehand.

## The effects of the CTF illustrated

It can be difficult to understand the major practical effects of the contrast transfer function in abstract. We therefore present a useful test object here, inspired by examples first presented in Downing and Glaeser (2008).

In the following figures, we mathematically apply the contrast transfer function to a test object. This test object is a wedge with horizontal stripes. The wedge gets narrower, and the stripes closer together, as we move from left to right. In this way, the wedge comprises a smooth range of (vertical) spatial frequencies, starting with the low frequencies at the left and finishing with the high frequencies at the right.

In these simulations, each pixel represents 1 Å. The wedge is therefore approximately 20 nm tall at its tallest side, and contains frequencies corresponding to the resolutions from approximately 26 Å to 6 Å. Additionally, we apply the contrast transfer function only in the vertical direction to prevent information from one spatial frequency spilling into adjacent columns.

<figure><img src="/files/0m2qvVdgHd5SQHdl3Z4b" alt="Top: a wedge with alternating black and white stripes. No CTF is applied to this wedge. On the left hand side, the stripes are wide, so their spatial frequency is low. On the right, the stripes are small, so the spatial frequency is high. Middle: the same wedge with a -1.5 micron defocus. The region in the middle of the wedge is flat grey, since there is no contrast at this frequency. Signal is delocalized away from its position in the wedge with no CTF. Bottom: a graph of the CTF with zero crossings marked."><figcaption></figcaption></figure>

At this defocus, the CTF has a negative value for low frequencies. This means that low frequencies (e.g., the left side of the wedge) have negative contrast.

Approximately halfway along the wedge, the CTF crosses zero (dashed vertical lines indicate zero-crossings). The frequencies at and near this point have no contrast — they are the same flat grey as the background. The zero-crossings of the CTF are why it is important to collect data at a range of defocus values. If all of these zero crossings were at the same point, that frequency would never have any contrast, and the images would therefore be incapable of producing a 3D volume.

After the first zero crossing, the CTF has a positive value. This means the wedge is again visible against the grey background, but black has become white and vice versa. This is another effect of the CTF that must be corrected for. In reality, the stripes have the same “density” all the way along the wedge. This flipping is purely an effect of the CTF.

The image above shows the CTF at only a single defocus (-1.5 µm). In the below animation, the defocus varies smoothly from 0.0 to 3.0 µm. The top pane shows the frequency wedge image while the bottom pane shows the CTF. Both are at the same defocus, and the frequencies are aligned in each.

<figure><img src="/files/DyRzlMcEMIKO2MvXfz3M" alt="An animation of the wedge from the previous image. Defocus increases from 0 to -3 microns."><figcaption></figcaption></figure>

Note first that as defocus increases, so does the contrast at the low frequencies (left end of the wedge). At the beginning of the animation, when the defocus is 0, the left side of the wedge is nearly invisible. As defocus increases, so does the contrast of the low-frequency, left-hand side of the wedge.

Next, observe that the zero crossings move up and down along the wedge as the defocus changes, since the contrast transfer function crosses zero at different frequencies depending on the defocus. Zero crossings are indicated in the bottom pane with dashed lines.

Finally, watch the high-frequency (right-hand) tail of the wedge as defocus changes. The information from this high-frequency region is displaced by a significant fraction of the total size of the object at high defocus!

## Useful Resources

For an approachable discussion of the contrast transfer function, we recommend Grant Jensen’s lecture on the topic, available [on YouTube](https://youtu.be/mPynoF2j6zc).

Readers interested in the math behind phase contrast and the contrast transfer function may find the [notes from Fred Sigworth and Hemant Tagare](https://cryoemprinciples.yale.edu/chapters) interesting. [Marin van Heel’s notes](https://www.singleparticles.org/methodology/MvH_Phase_Contrast.pdf) also provide interesting discussion of the topic. Finally, [Transmission Electron Microscopy by Kohl and Reimer](https://link.springer.com/book/10.1007/978-0-387-40093-8) provides a thorough and detailed reference for motivated readers.

To check whether CTF aliasing occurs for a given set of experimental parameters in 2D, consider using [this online script](https://3dem.github.io/relion/ctf.html), written by Takanori Nakane.

[This tool](https://ctfsimulation.streamlit.app/) from the Jiang lab simulates a CTF from a variety of user-selected parameters. It can be a helpful way to build intuition about the effects of various microscope parameters on the final CTF.

Another useful tool for developing intuition about the CTF’s effect on images is available from [Johannes Elferich](https://jojoelfe.github.io/webgl-ctf/image_abb).

## References

1. Downing, K. H. & Glaeser, R. M. Restoration of weak phase-contrast images recorded with a high degree of defocus: The “twin image” problem associated with CTF correction. *Ultramicroscopy* **108**, 921–928 (2008).

## CTF Estimation Jobs

{% content-ref url="/pages/-MR1qOsE3pv\_NChlhkcR" %}
[Job: Patch CTF Estimation](/processing-data/all-job-types-in-cryosparc/ctf-estimation/job-patch-ctf-estimation)
{% endcontent-ref %}

{% content-ref url="/pages/-MSTwb\_Lb0qdUOYUKrjD" %}
[Job: Patch CTF Extraction](/processing-data/all-job-types-in-cryosparc/ctf-estimation/job-patch-ctf-extraction)
{% endcontent-ref %}

{% content-ref url="/pages/ImRIxxhDrKW4PJmyJ3M0" %}
[Job: CTFFIND4 (Wrapper)](/processing-data/all-job-types-in-cryosparc/ctf-estimation/job-ctffind4-wrapper)
{% endcontent-ref %}

{% content-ref url="/pages/T9YbsSqKDdR42S8P0v9x" %}
[Job: Gctf (Wrapper) (Legacy)](/processing-data/all-job-types-in-cryosparc/ctf-estimation/job-gctf-wrapper-legacy)
{% endcontent-ref %}

## CTF Estimation Tutorials

{% content-ref url="/pages/-MNeA4BFZRumxWAGw8dk" %}
[Tutorial: Patch Motion and Patch CTF](/processing-data/tutorials-and-case-studies/tutorial-patch-motion-and-patch-ctf)
{% endcontent-ref %}


# Job: Patch CTF Estimation

Patch-based CTF estimation.

## At a Glance

Estimate the Contrast Transfer Function for micrographs, accounting for the fact that the sample may not be perfectly flat.

## Description

Patch-based CTF estimation automatically estimates a defocus landscape for tilted, bent, deformed samples and is accurate for all particle sizes and types including flexible and membrane proteins. This is accomplished by a fast GPU implementation that usually takes 1 - 2 seconds per micrograph. No prior knowledge about particle locations is needed.

## Inputs

### Micrographs

Patch CTF Estimation requires motion-corrected micrographs, typically straight from a [Patch Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction) job.

## Commonly Adjusted Parameters

The default parameters for this job are often sufficient for high-quality results.

## Outputs

### Micrographs processed

Micrographs with CTF estimates. During particle extraction, these estimates are used to calculate a local CTF at a particle’s precise location.

### Micrographs incomplete

Micrographs which were not successfully corrected are output in this group. Typically this is due to some error specific to the micrograph file — the Event Log should contain more information on why a particular micrograph failed.

### Diagnostic plots

Several diagnostic plots are also produced for each of a number of micrographs.

#### 1D search plot

<figure><img src="/files/CHTnyqxfcEeEtVZGxGij" alt="A plot of the quality of fit with various defoci. A clear peak is visible just over 2 microns."><figcaption></figcaption></figure>

This plot indicates the quality of fit for a 1D search over defocus. Higher values indicate a better fit between the idealized CTF at that defocus and the data. Ideally, a single sharp peak is observed, indicating that there is one defocus that matches much better than the others.

#### 2D defocus landscape

<figure><img src="/files/P140iafWouQM93BMviH6" alt="A plane is graphed in 3D space. The X and Y coordinates are in pixels, while the height of the surface (in Z) represents the defocus."><figcaption></figcaption></figure>

The defocus landscape displays the modeled defocus over the entire micrograph. Typically, this landscape should display some variation that captures the shape of the ice layer.

#### CTF Fit

The CTF fit helps you assess the quality of your data, and the quality of the CTF fit to that data. There is a lot of information in this plot, so we will build it up piece by piece.

First, the black line is the radial average of the power spectrum of the image. The Thon rings (used to fit the CTF) are present as oscillations between 0 and 1 in this image. This plot accounts for astigmatism using equi-phase averaging (e.g., similar to GCTF \[1]).

<figure><img src="/files/2apfPQ99Ynd8sSjCZ4yO" alt="A graph of the power spectrum of a micrograph. It oscillates between positive and negative numbers as the frequency increases."><figcaption></figcaption></figure>

The red line is the ideal CTF which has been found by Patch CTF Estimation. Note that this is the ideal CTF at the average defocus across the micrograph — each patch will have a slightly different function fitted to it. Ideally, this red line perfectly coincides with the black line.

<figure><img src="/files/iq56jWN3IZkC5E4Hvjf7" alt="The CTF fit to the previous power spectrum is plotted in red."><figcaption></figcaption></figure>

We can measure how well the power spectrum and the ideal CTF match by calculating the correlation between the two. This is plotted in blue. The point at which this line crosses 0.3 is typically called the CTF fit resolution, and is marked in the plot with a green vertical line. Note that this is not necessarily a hard limit on the quality of your data, and the images are not filtered to this resolution — it simply gives you an estimate of what quality to expect from this image.

<figure><img src="/files/iVj9uPLIlU9CqLF3R2M7" alt="A cyan line plots the correlation between the power spectrum and CTF"><figcaption></figcaption></figure>

The final image, with all three of these lines, is displayed below. Note that the plot title also contains some useful information

* The defoci on the major and minor axes are reported in Å (DF1 and DF2, respectively)
* The astigmatism angle, in radians, is given (ANGAST)
* The phase shift (if any) is reported (PHASE)
* The fit resolution is reported in Å (FIT)

<figure><img src="/files/kOSRokXedZ1DR7VesHSS" alt="The power spectrum, CTF, and correlation are plotted in black, red, and cyan respectively."><figcaption></figcaption></figure>

### Ice thickness

<figure><img src="/files/ViUwprlSNXu73LWAzS8T" alt="A plot with an unlabeled y-axis showing a peak around 1/4 Å. This peak is associated with amorphous ice and is used to approximate ice thickness."><figcaption></figcaption></figure>

The relative ice thickness is measured by comparing the background signal in a band centered on 0.265 Å-1 (indicated by a blue fill) to a wider band which includes this region (indicated by the green bars). If this band has high background, it means there is more ice in the image which, assuming ice takes up approximately the same proportion of each image, means the ice is thicker.

This plot does not directly report anything about data quality or CTF fit quality, although it is generally taken to be true that thinner ice results in higher quality reconstructions.

## Common Next Steps

Micrographs with CTF estimates are ready for particle extraction — a common workflow is performing [Blob Picking](/processing-data/all-job-types-in-cryosparc/particle-picking/job-blob-picker) and [Particle Extraction](/processing-data/all-job-types-in-cryosparc/extraction/job-extract-from-micrographs) at first, then using the initial results for more advanced particle picking jobs.

After a high-quality map is achieved, data may benefit from [CTF Refinement](/processing-data/tutorials-and-case-studies/tutorial-ctf-refinement) to improve defocus estimates and account for higher-order aberrations.

## Implementation Details

The overall implementation of Patch CTF follows a similar pathway to [Patch Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction). First, a coarse estimate of the CTF is modeled for the entire micrograph. Then, patches are used to estimate a function which returns the modeled defocus for a given micrograph coordinate.

For an initial, coarse estimate of the CTF, it is assumed that there is no astigmatism. A simple correlation with the radially-averaged power spectrum is used to find the best-fitting defocus.

A new envelope function is then calculated using this coarse defocus estimate, after which the 2D CTF is estimated for the entire micrograph including astigmatism. The estimated defocus is refined for each patch.

These patch CTF estimates are used to fit a spline function which provides the estimated defocus at a given (x, y) coordinate on the micrograph. Ultimately, particle defocus estimates come from this spline function. The `Override knots y` and `Override knots X` parameters control the degrees of freedom of this spline function. Note that changing the number of knots **does not** change the number of patches used by Patch CTF Estimation. See the [Patch Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction) job for a more detailed explanation of these parameters.

## References

1. Zhang, K. Gctf: Real-time CTF determination and correction. *J. Struct. Biol.* **193**, 1–12 (2016).
2. Tegunov, D. & Cramer, P. Real-time cryo-EM data pre-processing with Warp. *bioRxiv* (2018) doi:[10.1038/s41592-019-0580-y](https://doi.org/10.1038/s41592-019-0580-y).


# Job: Patch CTF Extraction

Patch CTF extraction.

## At a Glance

<figure><img src="/files/P31ZaakmYZNMjadTNEMm" alt="A flowchart. We start with micrographs without CTF estimates. Particle images are picked from these micrographs, and the micrographs are separately run through Patch CTF Estimation. Patch CTF Extraction is used to produce particle images with CTF estimates."><figcaption></figcaption></figure>

Apply CTF estimates from micrographs to particle locations which were picked or extracted before CTF estimation.

## Description

If particles are picked before CTF estimation is performed, those particles will not have CTF estimates and cannot be used in downstream analyses. Rather than re-pick or re-extract particles from micrographs with CTF estimates, Patch CTF Extraction allows particle CTF estimates to be updated directly.

In a typical workflow particles are picked after a [Patch CTF](/processing-data/all-job-types-in-cryosparc/ctf-estimation/job-patch-ctf-estimation) job. In that case this job is unnecessary, since the particles would have CTF estimates from their extraction job.

## Inputs

### Micrographs

This job requires micrographs with CTF estimates, typically from a [Patch CTF](/processing-data/all-job-types-in-cryosparc/ctf-estimation/job-patch-ctf-extraction) job.

### Particles

This job accepts either particle picks or extracted particle images. If picks are provided, they will have to be [extracted](/processing-data/all-job-types-in-cryosparc/extraction/job-extract-from-micrographs) before downstream use.

## Commonly Adjusted Parameters

### Flip mic. in x/y before extract

If particles were picked in external software, the coordinates may need to be flipped using one or both of these parameters before extraction. If CryoSPARC was used to pick particles, both of these parameters should be left off.

## Outputs

### Particles

If extracted particle images were provided as an input, the output particle images have a CTF and are ready for further use. If particle pick locations were provided, these particles need to be [extracted](/processing-data/all-job-types-in-cryosparc/extraction/job-extract-from-micrographs) before use.

## Recommended Alternatives

If providing particle pick locations (i.e., not extracted particles), it may be more convenient to use the Force re-extract CTFs from micrographs parameter of [Extract From Micrographs](/processing-data/all-job-types-in-cryosparc/extraction/job-extract-from-micrographs) rather than this job. This combines the CTF update and image extraction steps into a single job.


# Job: CTFFIND4 (Wrapper)

## At a Glance

Estimate the contrast transfer function parameters for exposures using CTFFIND4.

## Description

This job is a wrapper around CTFFIND4 (Rohou and Grigorieff 2015). Please read the license terms below.

Documentation and explanation of various parameters and plots are available from [the Grigorieff Lab](https://grigoriefflab.umassmed.edu/ctffind4).

### CTFFIND4 License Terms

> The Janelia Research Campus Software License 1.2 Copyright (c) 2018, Howard Hughes Medical Institute, All rights reserved. Redistribution and use in source and binary forms, with or without modification, are permitted provided that the following conditions are met: Redistributions of source code must retain the above copyright notice, this list of conditions and the following disclaimer. Redistributions in binary form must reproduce the above copyright notice, this list of conditions and the following disclaimer in the documentation and/or other materials provided with the distribution. Neither the name of the Howard Hughes Medical Institute nor the names of its contributors may be used to endorse or promote products derived from this software without specific prior written permission. THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS" AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, ANY IMPLIED WARRANTIES OF MERCHANTABILITY, NON-INFRINGEMENT, OR FITNESS FOR A PARTICULAR PURPOSE ARE DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR CONTRIBUTORS BE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES; LOSS OF USE, DATA, OR PROFITS; REASONABLE ROYALTIES; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.

## Inputs

### Exposures

Unlike CryoSPARC’s Patch CTF job, CTFFIND4 can process either movies or micrographs. If movies are provided, they will still need to be motion corrected before further processing can occur in CryoSPARC.

{% hint style="warning" %}
When using movies as input, ensure movie frames are gain corrected before import to CryoSPARC. **Movies connected to CTFFIND (Wrapper) which were imported into CryoSPARC with a separate gain correction file will not work**, as CryoSPARC does not bake gain correction into the movies during import. When you connect these movies to CTFFIND it will fail to recognize the gain correction and performance will likely suffer as a result.
{% endhint %}

## Commonly Adjusted Parameters

For convenience, the parameter tooltips reproduce help messages from CTFFIND4. For further assistance and suggested values, please consult the [CTFFIND4 documentation](https://grigoriefflab.umassmed.edu/ctffind4).

## Outputs

### Exposures

Exposures with CTF estimates. Exposures will be of the same type (i.e., micrographs or movies) as the inputs.

## Common Next Steps

If inputs were movies, [Patch Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction) is typically the next step. Otherwise, the micrographs are ready for particle picking and extraction.

## References

1. Rohou, A. & Grigorieff, N. CTFFIND4: Fast and accurate defocus estimation from electron micrographs. *Journal of Structural Biology* **192**, 216–221 (2015).


# Job: Gctf (Wrapper) (Legacy)

## At a Glance

Estimate the contrast transfer function parameters for exposures using Gctf.

## Description

This job is a wrapper around Gctf (Zhang 2016). See the original publication for an explanation of the algorithms employed by Gctf and recommendations for the diagnostic plots.

Note that this job is a Legacy job, and does not show up in the job builder by default.

## Inputs

### Exposures

Gctf can process either movies or micrographs. If movies are provided, they will still need to be motion corrected before further processing can occur in CryoSPARC.

### Particle stacks

Gctf optionally accepts particle stacks for local adjustment of CTF parameters. This process does not use a reference 3D structure. GCTF’s local refinement of CTF parameters is only expected to improve results that are already at relatively high resolution (e.g., better than 4 Å).

## Commonly Adjusted Parameters

### Abs path to Gctf executable

This dropdown selects from a list of Gctf binary names for various versions of CUDA. Note that these binaries **must** be installed at `deps/external/gctf-1.06/bin/` within the `cryosparc_worker` directory.

### Abs path to CUDA version

Publicly available Gctf binaries only support CUDA ≤ 8, but CryoSPARC uses a newer CUDA version. Users of this wrapper must therefore download and maintain their own copy of the CUDA 8 toolkit. The toolkit can be installed without root access by downloading the [CUDA 8 runfile from NVIDIA](https://developer.nvidia.com/cuda-80-ga2-download-archive) and installing it in a location of the user’s choice.

{% hint style="info" %}
As Gctf is compiled code, only the runtime toolkit is required. Carefully inspect the options while using the runfile and avoid installing any features other than the toolkit.
{% endhint %}

The absolute path of the `lib64` directory in this toolkit installation must be provided to this parameter.

## Outputs

### Exposures

Exposures with CTF estimates. Exposures will be of the same type (i.e., micrographs or movies) as the inputs.

## Common Next Steps

If inputs were movies, [Patch Motion Correction](/processing-data/all-job-types-in-cryosparc/motion-correction/job-patch-motion-correction) is typically the next step. Otherwise, the micrographs are ready for particle picking and extraction.

## References

1. Zhang, K. Gctf: Real-time CTF determination and correction. *J. Struct. Biol.* **193**, 1–12 (2016).


# Exposure Curation

Curate micrographs (exposures) to remove low-quality data.

## Exposure Curation Jobs

{% content-ref url="/pages/Q8E0OTFfuLwzoTaHWLu3" %}
[Job: Micrograph Denoiser (BETA)](/processing-data/all-job-types-in-cryosparc/exposure-curation/job-micrograph-denoiser-beta)
{% endcontent-ref %}

{% content-ref url="/pages/HabUzhCKiO7jmG5uXxaG" %}
[Job: Micrograph Junk Detector (BETA)](/processing-data/all-job-types-in-cryosparc/exposure-curation/job-micrograph-junk-detector-beta)
{% endcontent-ref %}

{% content-ref url="/pages/wEqFUe0gXE8bm8Ybl46q" %}
[Interactive Job: Manually Curate Exposures](/processing-data/all-job-types-in-cryosparc/exposure-curation/interactive-job-manually-curate-exposures)
{% endcontent-ref %}

## Exposure Curation Tutorials

{% content-ref url="/pages/-MNdjp0na854zYe\_BzMB" %}
[Tutorial: Manually Curate Exposures (v3)](/guides-for-v3/tutorial-manually-curate-exposures)
{% endcontent-ref %}


# Job: Micrograph Denoiser (BETA)

## At a Glance

<figure><img src="/files/2BaHaMKxxFrUYkgu9Ted" alt=""><figcaption><p>A micrograph from EMPIAR-10424 (Nakane et al. 2020)</p></figcaption></figure>

Produce enhanced micrograph images to aid particle picking and visual inspection.

## Description

Micrograph Denoiser takes micrographs as input and produces denoised versions of those micrographs. These denoised versions can be used downstream to help pick particles and to aid in visual inspection of the micrographs. Note that particles are extracted from the raw micrographs and not the denoised micrographs even if the latter are used for particle picking.

The denoiser works by first learning repeated patterns in the data during the training step. Then, to perform the denoising, it considers each region of the input micrograph. The denoiser is trained to match patterns it has learned and it amplifies those patterns in the denoised output micrograph. Features and noise which do not match are dimmed, thereby enhancing the visual quality of the denoised output micrograph. For a more thorough explanation of how the Micrograph Denoiser works, see the[ Denoiser Training section](#denoiser-training).

Denoised micrographs have significantly higher contrast than, for example, a lowpass filtered micrograph because image content is *added* by the denoiser. Of course, this contrast is not truly present in any single given region of the data — it is learned in aggregate by the denoiser. As such, the denoised images can be quite helpful for visual inspection or particle picking, but they do not contain increased signal for 2D or 3D reconstruction and therefore particles are not extracted from these denoised micrographs for downstream processing.

This job comprises both the training and application of the denoiser model. Input micrographs with training data are used to train the denoiser model, then *all* connected micrographs are denoised to produce the output.

## Inputs

### Exposures

Micrographs to be denoised, with background subtracted and CTF estimates available. The pixel size of these micrographs must be smaller than 3 Å.

{% hint style="info" %}
Currently, only Patch Motion Correction performs background subtraction, so the input micrographs must come from movies which were motion corrected by Patch Motion Correction.
{% endhint %}

If a new Denoising model will be trained (which is the recommended workflow), the input micrographs must also have training data. This data is only generated by Patch Motion Jobs run with CryoSPARC version **4.5 or later**. If you wish to perform denoising on data motion corrected in prior versions of CryoSPARC, see [Denoising Data from Existing Patch Motion Jobs](#denoising-data-from-existing-patch-motion-jobs).

Thus, a typical preprocessing workflow for CryoSPARC v4.5 or later might be

1. Import Movies
2. Patch Motion Correction
3. Patch CTF
4. Micrograph Denoiser
5. Blob Picker, etc…

### Denoise Model

If a Micrograph Denoiser has previously been run on this data, you may connect the Denoiser model output of the previous job to use that model instead of training a new one.

## Commonly Adjusted Parameters

### Renormalize input greyscale

In some situations, the denoiser must (or should) re-estimate the range of pixel values that likely correspond to particles (as opposed to empty ice or contaminants like crystalline ice or carbon). This estimation is called *greyscale normalization*, and `Renormalize input greyscale` controls whether or not this occurs. For more information on greyscale normalization, see [that section of this guide page](#greyscale).

When using the pre-trained model or when training a new model, this parameter is not displayed. This is because, in both of these cases, the model was not (or has not yet been) trained on the data and so must have a new greyscale normalization estimate.

When using a model trained by a previous Micrograph Denoiser job, this parameter becomes visible. If the input model was trained on the same data it will be denoising, it is typically not necessary to renormalize the greyscale and so this parameter can be kept off. If, however, it was trained on different data (even a different dataset of the same particle), it is likely worth renormalizing the greyscale and this parameter should be turned on.

### Greyscale normalization factor

Although the greyscale normalization procedure typically finds the correct range for the normalized greyscale, it is not perfect. This parameter (1.0 by default) is a multiplicative scale applied after the greyscale normalization process. For example, if the automated normalization determines that training should take place using a greyscale ranging from 0 to 200, but the `Greyscale normalization factor` is set to `1.5`, the final greyscale used during training and denoising will be 0 to 300.

Typically, this parameter can be left at the default value of 1.0. One notable exception is HexAuFoil grids, which often require a lower value for this parameter due to the significant fraction of the micrograph occupied by gold.

[The first plot produced by Micrograph Denoiser](#diagnostic-plots) is an example micrograph with the greyscale normalization applied. If this plot appears flat and grey or completely blown out, this parameter may need to be adjusted to a lower or higher value, respectively.

### Use pretrained model

If no input model is connected, you can either train a new model based on the input data or use the pre-trained model that is packaged with CryoSPARC. Generally, we recommend that this setting is left off, since a model trained directly from the data is expected to perform better and is relatively fast.

<figure><img src="/files/lztCV420JWPWYAO1ZZdb" alt=""><figcaption><p>The same micrograph (from EMPIAR 10335; Han et al. 2019) denoised using the pretrained model (left) or a model trained on this dataset (right). Training took eight minutes to complete 200 epochs with 100 training micrographs.</p></figcaption></figure>

### Number of mics for training

How many micrographs are used during denoiser training. Micrographs are selected randomly from the input for training, up to this number. If fewer than this number have training data available, the job will produce a warning but proceed with the training. At least 10 micrographs are required to train the denoiser; if fewer than 10 micrographs have training data, the job will fail.

We have generally found that 100 micrographs are sufficient to train a high-quality denoiser.

### Train from scratch

If this parameter is true (default), a new denoiser will be trained starting from a random initialization. This generally produces better results. If this parameter is false, the training will instead start with an initialization using the pre-trained denoiser as the starting point.

### Num training epochs

This parameter controls the number of times the denoiser is trained on the selected training subset. If the results of a denoiser job are still noisy or difficult to interpret, re-running the job with a greater number of epochs can produce better results at the cost of increased training time.

### Crop micrograph edges (fraction)

In some datasets, the edges of micrographs may have artifacts due to aberrations in the microscope or significant full-frame drift. These high-contrast artifacts can degrade denoiser training performance. In these cases, it is beneficial to ignore the edges of the micrograph during training.

This parameter sets a fraction of the micrograph to ignore *on each edge*. For example, setting `Crop micrograph edges` to 0.01 will crop 1 pixel off each side of a 100 x 100 pixel image, resulting in a 98 x 98 pixel training image. For rectangular images, the largest dimension is used to calculate the number of pixels: the same factor of 0.01 would trim a 50 x 100 pixel image to 48 x 98 pixels during training.

## Outputs

### Denoised micrographs

Denoised micrographs are output in this slot. Note that the original, non-denoised micrographs are also included in this same output. If this output is ultimately connected to an Extract Particles job, the non-denoised micrographs will be used automatically.

### Diagnostic plots

#### Normalization

The first plot produced by the Micrograph Denoiser is an example micrograph with normalization applied. See the [greyscale section of this page](#greyscale) for more information on how and why input micrographs are normalized before training and denoising. If this micrograph has little to no contrast, the normalization factor must be reduced. If this micrograph appears blown-out, with most pixels either white or black, the normalization factor must be increased.

<div data-full-width="true"><figure><img src="/files/1YcYL60tO8mDvjFnXKX4" alt=""><figcaption></figcaption></figure></div>

#### Training and Validation

Micrograph Denoiser also produces a plot of training and validation curves. However, unlike some machine learning tools, the weight of the various components of the Denoiser model change over the course of the training. We generally do not expect these plots to be informative, and users should focus on the denoised results to assess training quality.

<figure><img src="/files/xIJi5qixs4AhYk0nVjhM" alt=""><figcaption></figcaption></figure>

## Common Problems

### Denoised micrographs are blurry or still noisy

<figure><img src="/files/FS4l6Hp493J82DaHKMTn" alt=""><figcaption></figcaption></figure>

If the denoiser model has not yet converged, it will not be able to accurately model noise in the micrographs. This will result in blurry or still-noisy results. Increasing `Num training epochs` often helps in these cases.

### Denoiser produces images of empty ice or large, blotchy images

Flat, grey images or blotchy, high-contrast images, like those shown in the second row of the Diagnostic Plots above, are typically due to a Greyscale normalization factor that is too high or low, respectively. Changing this parameter should improve results.

## Common Next Steps

Denoised micrographs are especially helpful when performing and evaluating particle picking. Thus, a typical next step would be [Blob Picking](/processing-data/all-job-types-in-cryosparc/particle-picking/job-blob-picker) or [Template Picking](/processing-data/all-job-types-in-cryosparc/particle-picking/job-template-picker), followed by [Inspect Picks](/processing-data/all-job-types-in-cryosparc/particle-picking/interactive-job-inspect-particle-picks). The Inspect Picks job allows users to toggle between viewing the raw and denoised micrographs to evaluate pick locations. In CryoSPARC v4.6+, Inspect Picks is able to automatically cluster and select particles when picking is done on denoised micrographs (see [Interactive Jobs](/application-guide/interactive-jobs#interactive-job-inspect-particle-picks)).

After particles are picked, micrographs can be plugged directly into [Extract Micrographs](/processing-data/all-job-types-in-cryosparc/extraction/job-extract-from-micrographs) — the raw micrographs will automatically be used during extraction even if denoised micrograph images are available or were used for picking.

{% hint style="warning" %}
At this time, TOPAZ does not perform well on micrographs denoised with the Micrograph Denoiser. If particle picking will be performed with TOPAZ, we recommend using [TOPAZ Denoise](/processing-data/all-job-types-in-cryosparc/deep-picking/topaz/job-topaz-denoise-beta) instead.
{% endhint %}

## Denoising Data from Existing Patch Motion Jobs

Patch Motion Correction jobs run in versions of CryoSPARC prior to 4.5 do not generate the data necessary to train a denoiser model. To denoise these micrographs, two workflows are available:

1. *(recommended)* Performing Patch Motion Correction on a small subset of the movies to generate the necessary data, or
2. Using the pre-trained model

### Generate new training data (recommended)

In our testing, a denoiser trained on the input data typically outperforms the pre-trained model. We therefore recommend that the necessary training data is generated for a subset of micrographs and used to create the denoising model. This model can then denoise the entire set of micrographs, including those for which were not re-motion corrected.

1. Create a Patch Motion Correction job with the same settings as the existing job, except set `Only process this many movies` and `Num. movies for denoiser training data` both to `100`. Plug the micrographs from the initial Patch CTF Estimation job into the Exposures input.
2. Run the resulting micrographs through a Micrograph Denoiser job. The CTF estimates from the initial job will be used along with the training data from the new Patch Motion Correction job.
3. Set up a new Micrograph Denoiser job with
   1. the original, full set of motion-corrected and CTF-estimated movies, and
   2. the denoise model from the first Micrograph Denoiser job

This takes advantage of the benefits of training the denoiser on the data, while avoiding the need to re-motion correct the entire dataset.

### Use the pretrained method

If time is at a premium, plugging the existing movies into the Micrograph Denoiser and turning `Use pretrained model` on skips model training and uses the pretrained model. The results with the pretrained model are typically not as good as those trained on the data, but will likely still represent an improvement over simple lowpass filtering.

## Denoiser Training

The Micrograph Denoiser in CryoSPARC uses a specialized neural network architecture to produce denoised micrographs from training data. The neural network is trained using a *Nosie2Noise* methodology (Lehtinen et al. 2018) to predict what parts of an image are noise and what parts are signal.

The basic principle behind Noise2Noise training is very similar to that of GSFSC validation. First, training data is generated by splitting each movie into odd and even frames, then creating half-micrographs from only those frames. Any signal in the movie should be the same in both half-micrographs, but the noise is treated as entirely random and independent.

<div data-full-width="true"><figure><img src="/files/y7amT8okAvsmgmqgPlnJ" alt=""><figcaption></figcaption></figure></div>

Next, the neural network is trained to predict half-micrograph B from half-micrograph A. The only information that is the same between the two micrographs is the signal. Thus, as this neural network improves its ability to predict half B from half A, it is in effect learning patterns present only in the signal — modeling noise would not improve its ability to predict half B.

<div data-full-width="true"><figure><img src="/files/wBaYdnKALQTFze2e87rO" alt=""><figcaption></figcaption></figure></div>

This setup has been explored in several denoiser methods for cryo-EM data, including Warp (Tegunov and Cramer 2018), TOPAZ denoise (Bepler et al. 2020), and Sphire (Wagner et al. 2020).

In addition to the *Noise2Noise* training setup, CryoSPARC’s Micrograph Denoiser pre-corrects for the CTF in training data and input data, causing the denoised micrographs to be as visually consistent as possible across a range of defocus values. Furthermore, traditional denoising metrics are used to augment the training objective function to encourage the denoiser to quickly learn to produce visually clear micrographs emphasizing repeated signal such as particles.

## Greyscale

Each pixel in a cryoEM micrograph contains a numeric value representing the electron dose received at that pixel. It can have, essentially, any value. Typically, these numbers are represented by making the highest value pure black and the lowest value pure white, with intermediate values linearly scaled to an intermediate grey. This mapping from values to darkness is called the “greyscale” of the image. Each dataset will have a slightly different greyscale, depending on the electron dose, pixel size, aperture settings, etc.

Particles typically fall within a relatively narrow band of values across a dataset. What’s more, they are typically much closer in value to empty ice than to very dark objects like carbon or crystalline ice. Thus, if we trained the denoiser on the raw greyscale it may not even be able to detect true particles, since they would have essentially the same value as empty ice.

<figure><img src="/files/YDvp5oKN7Pq2jRX64zRe" alt=""><figcaption></figcaption></figure>

To avoid this flattening effect, the Micrograph Denoiser first estimates the greyscale band most likely to contain particles and produces a *normalized greyscale* covering only that range. Any values outside this range are clipped to black or white. This dedicates the greatest dynamic range to the values most likely to contain particles.

<figure><img src="/files/PfxPSEHBgkRnFCRWZo1Z" alt="" width="375"><figcaption></figcaption></figure>

Note in the figure above that, after normalization, the values corresponding to empty ice are all white, and any values outside the expected range for a particle are all black, regardless of their true value. All the variance of the greyscale is focused on the region in which particles are expected to lie. Normalizing the greyscale in this way focuses the denoiser on detecting patterns from the particles rather than empty ice or contaminants.

`Greyscale normalization factor` adjusts the estimated greyscale by multiplying the limits, moving them further from the mean. For example if, in the initial greyscale, any values below -10 are pure white and any values above 10 are pure black, then setting the `normalization factor` to `1.5` would result in a final greyscale (used for training) in which -15 and below is white and 15 and above is black.

<figure><img src="/files/QNp1xJhwKyeTr0tajhlj" alt="" width="374"><figcaption></figcaption></figure>

This means the model would be trained with a wider range of values; this may improve or degrade performance, depending on whether or not those values are useful for learning about the particles in the dataset.

## References

1. Nakane, T. *et al.* Single-particle cryo-EM at atomic resolution. *Nature* **587**, 152–156 (2020).
2. Han, Y. *et al.* High-yield monolayer graphene grids for near-atomic resolution cryoelectron microscopy. *Proceedings of the National Academy of Sciences* **117**, 1009–1014 (2019).
3. Lehtinen, J. *et al.* Noise2Noise: Learning Image Restoration without Clean Data. *arXiv* (2018) doi:[10.48550/arXiv.1803.04189](https://doi.org/10.48550/arXiv.1803.04189).
4. Tegunov, D. & Cramer, P. Real-time cryo-EM data pre-processing with Warp. *bioRxiv* (2018) doi:[10.1038/s41592-019-0580-y](https://doi.org/10.1038/s41592-019-0580-y).
5. Bepler, T., Kelley, K., Noble, A. J. & Berger, B. Topaz-Denoise: general deep denoising models for cryoEM and cryoET. *Nature Communications* **11**, 5208 (2020).
6. Wagner, T. & Raunser, S. The evolution of SPHIRE-crYOLO particle picking and its application in automated cryo-EM processing workflows. *Communications Biology* **3**, 61 (2020).




---

[Next Page](/llms-full.txt/1)

