============ Installation ============ Prerequisites ------------- Before building Cocoa, ensure you have the following dependencies installed: Compilers ^^^^^^^^^ Cocoa requires a C++20 compatible compiler. The minimum versions are dictated by Kokkos 5.0 (included with Trilinos 17), which sets stricter requirements than C++20 alone. .. list-table:: Minimum Compiler Requirements :header-rows: 1 :widths: 30 20 50 * - Compiler - Minimum Version - Notes * - GCC - 10.4.0 - Recommended for CPU and as CUDA host compiler * - Clang (CPU) - 14.0.0 - For CPU-only builds * - Clang (CUDA host) - 15.0.0 - When used as nvcc host compiler * - NVIDIA nvcc - 12.2 - Requires CUDA Toolkit 12.2+ * - Intel icpx (CPU) - 2022.0.0 - Intel oneAPI DPC++/C++ Compiler * - Intel icpx (SYCL) - 2024.2.1 - For SYCL backend builds * - ROCm (HIPCC) - 6.2.0 - For AMD GPU builds * - NVIDIA HPC SDK (NVC++) - 22.3 - Alternative to nvcc for NVIDIA GPUs .. note:: See the `Kokkos Requirements `_ documentation for the most up-to-date information. Build System ^^^^^^^^^^^^ - CMake 3.23 or later - GNU Make or Ninja build system Required Libraries ^^^^^^^^^^^^^^^^^^ The following libraries must be pre-installed on your system: - **Trilinos 17.0 or later** (with Kokkos, KokkosKernels, Tpetra, Belos, Ifpack2, Zoltan2 enabled) - **NetCDF-C** (4.9.3+ recommended; for mesh and output I/O) - **HDF5** (development headers required; installed automatically as a NetCDF-C dependency) .. warning:: NetCDF-C versions prior to 4.9.3 have a bug (`#2674 `_) that causes spurious HDF5 error messages on stderr when reading variables. Ubuntu 24.04 ships NetCDF-C 4.9.2; if using that distribution, build NetCDF-C 4.9.3+ from source. - **ParMETIS** (for mesh partitioning, required via Zoltan2 for MPI builds) .. note:: Trilinos 17.0+ is required because it ships Kokkos 5.0+, whose APIs Cocoa depends on. Automatically Fetched Dependencies ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ The following dependencies are automatically downloaded and built via `CPM `_ during CMake configuration: - yaml-cpp (configuration file parsing) - spdlog (logging) - fmt (string formatting) - Catch2 (unit testing) Optional Dependencies ^^^^^^^^^^^^^^^^^^^^^ - **CUDA Toolkit** (for NVIDIA GPU support, required for Trilinos CUDA build) - **ROCm** (for AMD GPU support, required for Trilinos HIP build) - **MPI** (for distributed computing, if Trilinos was built with MPI) Building from Source -------------------- Clone the Repository ^^^^^^^^^^^^^^^^^^^^ .. code-block:: bash git clone https://github.com/cocoaorg/cocoa.git cd cocoa Configure with a Preset ^^^^^^^^^^^^^^^^^^^^^^^ ``CMakePresets.json`` at the repository root holds the configurations CI builds plus two developer ones, so a build no longer needs a wall of ``-D`` flags. Presets locate Trilinos and NetCDF through the ``TRILINOS_DIR`` and ``NETCDF_DIR`` environment variables; export them once for your machine: .. code-block:: bash export TRILINOS_DIR=/path/to/trilinos/lib/cmake/Trilinos export NETCDF_DIR=/path/to/netcdf-c cmake --preset release cmake --build --preset release ctest --preset release ``cmake --list-presets`` lists them; ``cmake --build --list-presets`` and ``ctest --list-presets`` do the same for the build and test presets of the same names. .. list-table:: :header-rows: 1 :widths: 22 78 * - Preset - Purpose * - ``release`` - Optimized build, for running the model. * - ``dev`` - Debug build with the unit tests. ``cocoa_BACKEND`` auto-selects, so a GPU box builds for its GPU. GRIB is off because it needs an ecCodes install. * - ``dev-maintainer`` - Release build with ``cocoa_MAINTAINER_MODE``, the integration tests and the benchmarks -- the configuration to run before pushing. * - ``ci`` - What the serial CI job configures, GRIB included. * - ``ci-cuda`` - The CUDA compile check. Its ``Trilinos_DIR`` names the CI image's CUDA Trilinos install; pass ``-DTrilinos_DIR=...`` to point it elsewhere. * - ``ci-tidy`` - Compile database for the clang-tidy gate; configures into ``build-tidy`` and is never built. * - ``ci-coverage`` - The instrumented build behind the coverage job. Every preset but ``ci-tidy`` configures into ``build/``, so switching between them reconfigures that one tree rather than creating a second one. Pass ``-B`` to put a preset somewhere else. A ``-D`` on the command line overrides the preset, so ``cmake --preset release -Dcocoa_ENABLE_GRIB=ON`` is enough for a one-off change. For a permanent one, put a ``CMakeUserPresets.json`` beside ``CMakePresets.json`` (it is untracked) and inherit: .. code-block:: json { "version": 4, "configurePresets": [ { "name": "my-release", "inherits": "release", "cacheVariables": { "Trilinos_DIR": "/home/me/trilinos/lib64/cmake/Trilinos", "NETCDF_DIR": "/home/me/spack/opt/spack/netcdf-c-4.9.3" } } ] } Configure with CMake ^^^^^^^^^^^^^^^^^^^^ Presets are a convenience, not a requirement; every option below can be passed directly. **Basic Build**: .. code-block:: bash mkdir build && cd build cmake .. \ -DCMAKE_BUILD_TYPE=Release \ -DNETCDF_DIR=/path/to/netcdf-c \ -DTrilinos_DIR=/path/to/trilinos CMake Options ^^^^^^^^^^^^^ .. list-table:: :header-rows: 1 :widths: 28 47 25 * - Option - Description - Default * - ``NETCDF_DIR`` - Hint path for CMake's ``FindNetCDF`` module. Point to the NetCDF-C installation prefix. - (auto-detected; required if not in system paths) * - ``Trilinos_DIR`` - Path to the Trilinos CMake config directory (e.g., ``/lib/cmake/Trilinos``). - (auto-detected; required if not in system paths) * - ``cocoa_BACKEND`` - Combined execution space and MPI configuration. Available options depend on the Trilinos build. Examples: ``CUDA+MPI``, ``CUDA``, ``HIP+MPI``, ``OPENMP+MPI``, ``OPENMP``, ``SERIAL+MPI``, ``SERIAL``. - ``DEFAULT`` -- auto-selects the best available backend from Trilinos (prefers GPU over CPU, MPI over non-MPI) * - ``CMAKE_BUILD_TYPE`` - Build type. Options: ``Release``, ``Debug``, ``RelWithDebInfo``, ``MinSizeRel``. - ``RelWithDebInfo`` (if not specified) * - ``CMAKE_INSTALL_PREFIX`` - Installation directory for ``cmake --install``. - ``/usr/local`` * - ``BUILD_TESTING`` - Build the unit test suite (requires Catch2, fetched automatically). - ``ON`` when Cocoa is the top-level project, ``OFF`` when vendored * - ``cocoa_MAINTAINER_MODE`` - Enable strict compiler warnings, sanitizers, cppcheck, and hardening. Must be requested explicitly; developer builds and CI pass it. - ``OFF`` * - ``cocoa_CUDA_MEMORY_SPACE`` - CUDA memory space. ``CUDA`` for device memory, ``CUDAUVM`` for unified virtual memory. Only applies to CUDA backends. - ``CUDA`` * - ``cocoa_USE_THRUST`` - Use ``thrust::copy_if`` for the wet/dry stream compaction on OpenMP and Serial backends (advanced). See :ref:`stream-compaction` below. Requires a Thrust/CCCL installation. CUDA builds always use CUB and ignore this option. - ``OFF`` * - ``cocoa_ENABLE_GRIB`` - Enable GRIB2 meteorological forcing (GFS, HRRR, ...). Requires an installed ECMWF ecCodes (found via ``eccodes_DIR``) built with the JPEG2000 (Jasper or OpenJPEG) and CCSDS/AEC codecs that NCEP products use; cocoa checks for those features at configure time. See :doc:`/user_guide/meteorological_forcing`. - ``OFF`` * - ``eccodes_DIR`` - Path to the ecCodes CMake config directory (e.g., ``/lib/cmake/eccodes``). Required with ``cocoa_ENABLE_GRIB=ON`` if ecCodes is not in the default search paths. Install a suitable build with, e.g., ``spack install eccodes +memfs +aec jp2k=openjpeg``. - (auto-detected) * - ``cocoa_PRINT_LOG_TIME`` - Print elapsed wall-clock time in screen log output. - ``OFF`` Floating-Point Precision ^^^^^^^^^^^^^^^^^^^^^^^^ Cocoa computes in double precision (FP64). There is no build-time precision option: to reduce GPU memory traffic, a fixed set of bandwidth-sensitive fields is *stored* as float and promoted to double on read, while all arithmetic stays in double. See :doc:`/theory/numerical_methods` for the list of mixed-precision fields and the rationale. .. _stream-compaction: Stream Compaction ^^^^^^^^^^^^^^^^^ Each step, the wet/dry solver rebuilds compacted lists of the currently wet elements and nodes -- a stream compaction (``copy_if`` over an index range). The implementation is selected by backend: - **CUDA**: ``cub::DeviceSelect::If`` on the Kokkos execution-space stream, with the algorithm's temporary storage allocated once at setup and reused every step (no per-step ``cudaMalloc``/``cudaFree``). CUB ships with the CUDA Toolkit, so no extra setup is needed. - **OpenMP/Serial with** ``cocoa_USE_THRUST=ON``: ``thrust::copy_if`` through the matching thrust host device system. - **Otherwise**: a fused Kokkos ``parallel_scan``. All implementations use the same selection predicate and produce identical results. Enabling ``cocoa_USE_THRUST`` requires a `Thrust `_ installation, typically via `NVIDIA CCCL `_, with CMake pointed at it: .. code-block:: bash cmake .. \ -Dcocoa_BACKEND=OPENMP+MPI \ -Dcocoa_USE_THRUST=ON \ -DCCCL_DIR=/path/to/cccl/lib/cmake/cccl # or, for a standalone Thrust: # -DThrust_DIR=/path/to/thrust/lib/cmake/thrust If ``cocoa_USE_THRUST=ON`` is requested but no Thrust/CCCL installation is found, or the backend is neither OpenMP nor Serial (and not CUDA, where the option is ignored), configuration fails with an explanatory error. Compile ^^^^^^^ .. code-block:: bash cmake --build --preset release # or, from the build directory, make -j$(nproc) Install ^^^^^^^ .. code-block:: bash cmake --install build Verifying the Installation -------------------------- Run the test suite to verify your installation: .. code-block:: bash ctest --preset release # or, from the build directory, ctest --output-on-failure The test presets exclude the ``validation`` label, which covers the channel studies that compare against analytic solutions rather than committed references. Run those deliberately with ``ctest -L validation``.