Aller au contenu

Scaling open-source industrial solvers on Cloud HPC: Performance benchmarks and configuration guidelines for code_saturne and code_aster on Oracle Cloud Infrastructure

Scaling open-source industrial solvers on Cloud HPC: Performance benchmarks and configuration guidelines for code_saturne and code_aster on Oracle Cloud Infrastructure

Accelerating CFD simulations in the Cloud: Challenges and opportunities

In modern engineering workflows, Computational Fluid Dynamics (CFD) plays a critical role in the design and validation of complex systems.

However, as simulation fidelity increases, so do computational demands.

Engineers are now routinely faced with:

  • large-scale meshes involving millions of cells,
  • long simulation times driven by complex physics models (e.g., LES turbulence),
  • and the need to run multiple design iterations within tight project timelines.

To address these challenges, High-Performance Computing (HPC) has become essential. By distributing computations across many processors, simulation time can be significantly reduced.

Cloud platforms now play a key role in this evolution by providing access to scalable computing resources without the constraints of on-premises infrastructure. Oracle Cloud Infrastructure (OCI) provides HPC environments specifically engineered for tightly coupled workloads such as industrial simulation, combining high core density, low-latency interconnects, and predictable performance.

However, achieving efficient performance in these environments is often convoluted.

While increasing the number of cores theoretically reduce computation time, real-world CFD applications often exhibit:

  • diminishing returns beyond a certain level of parallelism,
  • sensitivity to hardware architecture,
  • and strong dependence on Message Passing Interface (MPI) configuration and process placement.

As a result, simply allocating more computational resources does not always lead to faster simulations.

This reality leads to a key question for engineering teams adopting cloud HPC:

How to configure CFD solvers to efficiently leverage available HPC resources?

Open-Source CFD: Combining industrial robustness and scalability

Achieving efficient performance in HPC and cloud environments is not just a matter of adding more cores. Although simulation solvers are inherently parallel, their actual efficiency depends strongly on how computations are distributed, how memory is accessed, and how communication between processes is handled.

This is where HPC expertise becomes critical, regardless of whether the solver is closed or open-source. Proper configuration, process placement, and understanding of hardware architecture are essential to achieve optimal performance and control simulation costs.

In this context, open-source simulation tools offer an additional advantage.

By removing "per-core" or "per token" license fee constraints, they allow engineers to run large numbers of simulations in parallel and scale computations freely, up to their needs.

However, their adoption in industrial environments raises legitimate questions whether for CFD, structural mechanics, or other simulation disciplines. Engineers need to ensure that open-source solutions provide the same level of modeling capabilities, validation, and performance as established by close-source tools.

In this regard, the open-source simulation codes developed by EDF offer a compelling option.

Among these tools, code_saturne stands out as a mature CFD solver designed for industrial applications. Developed and maintained by EDF, it supports advanced turbulence modeling, complex geometries, and large-scale simulations. Released under open-source licenses, it also benefits from contributions and feedback from a broader community of industrial and academic users.

Overall, EDF's open-source suite (including code_aster, code_saturne, and openTELEMAC) was developed to address thought engineering challenges in highly regulated sectors such as energy and infrastructure. These codes integrate decades of R&D, are subject to rigorous validation processes, and are used daily in production environments by EDF and by a growing community of industrial users worldwide.

As a result, open-source simulation tools now combine robustness, versatility, and industrial credibility, making them strong candidates for companies looking to move away from closed-source solutions.

Yet, as discussed above, unlocking the full performance potential of these tools on cloud platforms requires careful attention to HPC configuration, an area where Simvia brings dedicated expertise.

At Simvia, our approach focuses on bridging this gap by enabling efficient deployment of EDF open-source solvers on cloud infrastructures and ensuring that simulations are properly configured to achieve optimal performance.

To illustrate this, we conducted a benchmark study using code_saturne on Oracle Cloud Infrastructure, based on a typical industrial CFD case.

Benchmarking CFD performance on Oracle Cloud

To evaluate the performance of open-source CFD simulations in a cloud environment, a benchmark study was conducted using code_saturne on Oracle Cloud Infrastructure (OCI).

The goal was to assess how a typical industrial CFD case scales with increasing computational resources, and to identify the key factors influencing performance.

A representative industrial CFD case

Figure 1: View of the partitioned volume mesh used for benchmark.
Figure 1: View of the partitioned volume mesh used for benchmark.

The selected test case is based on a steady-state RANS simulation, archetypal of industrial workflows, and more specifically of flow configurations encountered in nuclear fuel rod bundles.

The model consists of a 12.9 million cell mesh combined with a standard RANS turbulence modeling approach, reflecting common industrial CFD practices.

Cloud infrastructure and hardware configuration

Simulations were conducted on Oracle Cloud Infrastructure using AMD EPYC-based "BM.Standard.E5.192" bare metal instances, exemplary of modern HPC cloud environments. Each compute node is powered by AMD EPYC 9J14 processors (base frequency of 2.4 GHz, up to 3.7 GHz boost) and provides 192 physical cores across two sockets, along with 2304 GB of memory. The nodes deliver up to 100 Gbps of network bandwidth and run a Linux operating system.

The study focuses on strong scaling performance, analyzing how computation time evolves as the number of MPI processes increases for a fixed problem size.

Performance metrics and methodology

The main metric used in this study is the total simulation runtime, evaluated as a function of the number of MPI processes.

Several configurations were explored to assess their impact on performance, including variations in the number of MPI processes and process placement strategies (binding and mapping).

This approach makes it possible to identify scaling limits, understand performance bottlenecks, and evaluate the sensitivity of the solver to HPC configuration choices.

A practical objective: understanding real-world scaling behavior

Rather than focusing on theoretical peak performance, this benchmark aims to reflect real-world usage conditions.

In practice, engineers are faced with a simple question: how many cores should be used to achieve the best performance-to-cost ratio?

By analyzing the evolution of runtime across a range of configurations, this study provides tangible insights into the optimal operating range of the solver, the onset of performance saturation, and the impact of HPC configuration on simulation efficiency.

Understanding scaling behavior on Oracle Cloud

The results are presented in Figure 2, which shows the evolution of simulation runtime as a function of the number of MPI processes for different MPI configurations.

The dashed black curve represents the ideal scaling, while the colored curves correspond to different process placement strategies.

A key takeaway from these results is that scaling behavior is not intrinsic to the solver, but strongly depends on how it is configured.

This is why modern cloud platforms such as OCI are particularly valuable, enabling rapid exploration of multiple configurations without dedicated infrastructure provisioning. Access to recent hardware generations also allows performance evaluation under conditions representative of current HPC systems. This is clearly illustrated by the results obtained in this study. Among the tested configurations, the unconstrained placement (--bind-to none) consistently delivers the best performance. In this case, the operating system dynamically distributes MPI processes across available cores, enabling efficient use of memory bandwidth and overall hardware resources.

This configuration achieves strong scalability up to approximately 32 to 48 MPI processes, with performance remaining close to the ideal scaling. In this range, computational resources are effectively utilized, providing the best balance between time-to-solution and resource usage.

Beyond this point, however, performance gains show diminishing results.

For the present case, this corresponds to a ratio of about 250 000 cells per MPI process, providing a practical guideline for achieving efficient parallel performance. This type of rule-of-thumb is particularly valuable when sizing cloud resources for CFD workloads.

In contrast, other configurations show significantly less favorable scaling behavior. When processes are explicitly bound to cores (--bind-to core --map-by core), MPI processes are first concentrated on a single socket, leading to early memory bandwidth saturation. The second socket is only utilized at higher process counts, resulting in delayed performance improvement.

A more balanced distribution across sockets (--map-by ppr) improves resource utilization and leads to more stable performance compared to simple core-based mapping. However, it still does not outperform the unconstrained configuration, highlighting that manual binding strategies do not always lead to optimal results.

Beyond approximately 48 MPI processes, even in the best configuration, the results show a clear saturation effect, with runtime stabilizing despite the increase in computational resources.

This behavior reflects a transition toward a memory-bound regime, where memory bandwidth and inter-process communication become the dominant limiting factors.

Turning insights into practical guidelines

From an operational perspective, these results provide clear guidance for engineering teams.

Rather than simply increasing the number of cores, it is essential to identify the optimal operating range of the simulation. In this case, using more than 250 000 cells per MPI process leads to diminishing returns, with limited impact on overall performance.

=> Efficient HPC usage is therefore not about maximizing resources, but about using them wisely.

From benchmark to operational performance

This study highlights a key takeaway for engineering teams leveraging cloud HPC: achieving optimal performance is not only about accessing powerful hardware, but about understanding how simulation codes behave in parallel environments.

Even for a standard CFD case, performance is strongly influenced by the number of MPI processes, process placement strategies, and the balance between computation, memory access, and communication.

The results show that efficient scaling can be achieved within a well-defined operating range, beyond which additional resources provide diminishing returns. More importantly, this study demonstrates that relatively simple configuration choices can have a major impact on performance, making HPC expertise a key factor in unlocking the full potential of cloud-based simulations.

Figure 2: Simulation runtime versus number of MPI processes on Oracle Cloud Infrastructure for different MPI configurations.
Figure 2: Simulation runtime versus number of MPI processes on Oracle Cloud Infrastructure for different MPI configurations.

Bridging CFD expertise and Cloud HPC

At Simvia, our objective is precisely to bridge the gap between advanced simulation tools and modern HPC infrastructures.

By combining deep knowledge of EDF open-source solvers with expertise in HPC environments and cloud platforms, we help engineering teams deploy simulations efficiently, understand scaling behavior, and optimize performance for their specific use cases.

This approach enables users to fully benefit from the flexibility of cloud computing while maintaining control over simulation costs and turnaround times.

Beyond CFD: Extending to Structural Mechanics

While this study focuses on CFD with code_saturne, the same challenges (and opportunities) apply to other types of simulations.

Code_aster, EDF's open-source finite element solver for structural mechanics, addresses a wide range of industrial use cases, including non-linear structural analysis, fatigue and fracture mechanics, and large-scale multi-physics simulations.

Like code_saturne, code_aster is designed for high-performance computing and follows similar scaling principles, where efficient parallelization and appropriate configuration are essential to fully leverage available resources.

Towards scalable, open, and efficient simulation workflows

By combining robust open-source solvers such as code_saturne and code_aster with optimized deployment on cloud HPC platforms, engineering teams can accelerate simulation workflows, reduce computational costs, and increase their ability to explore design spaces.

  • With the right expertise, cloud computing becomes not just a resource, but a true accelerator for engineering performance.

Extending the benchmark perspective to structural mechanics with code_aster

While the present study primarily focuses on CFD workloads with code_saturne, the same performance and sizing questions also arise in structural mechanics. Large-scale finite element simulations involve their own computational challenges, especially when non-linear material behavior, large meshes, and sparse linear system solves are combined within iterative solution procedures.

In this context, code_aster, EDF's open-source finite element solver for structural mechanics, provides a relevant complementary perspective. Widely used for industrial-grade structural analysis, code_aster addresses a broad range of applications, including linear and non-linear mechanics, thermo-mechanical coupling, contact, fracture, and multi-physics simulations. As with CFD solvers, achieving efficient time-to-solution on cloud HPC infrastructure is not only a matter of hardware provisioning, but also of choosing the appropriate numerical strategy and parallel configuration. This naturally extends the broader message of the paper: cloud performance is unlocked through the combination of robust open-source solvers, adapted numerical methods, and HPC expertise.

A representative code_aster benchmark on Oracle Cloud Infrastructure

To illustrate these considerations on the structural mechanics side, a benchmark was conducted with code_aster on Oracle Cloud Infrastructure Bare-Metal instance using an AMD-based "BM.Standard.E5.192" shape. The selected test case is a quasi-static simulation of a turbine blade involving elasto-plastic material behavior. Such a case is typical of demanding mechanical analyses in which both non-linearity and the cost of linear algebra play a central role in overall performance.

Blade engine mesh and partitioning using PTSCOTCH
Blade engine mesh and partitioning using PTSCOTCH

From a numerical standpoint, the study was initially based on the MUMPS sparse direct solver, which is a well-established reference in large-scale finite element analysis. Additional tests were then carried out using PETSc-based strategies, in particular a GMRES iterative solver combined with a MUMPS-based preconditioning strategy in single precision. This comparison makes it possible to go beyond a pure hardware benchmark and assess how solver choices interact with cloud HPC resources.

Strong scaling behavior and solver comparison

The results, charted on figure 3, highlight a point that is well known to HPC practitioners but often underestimated in operational simulation workflows: for large structural mechanics problems, performance is driven not only by the number of cores, but also by the algebraic solution strategy.

Figure 3: Time-to-solution versus number of cores for the code_aster turbine blade benchmark on Oracle Cloud Infrastructure, comparing the MUMPS direct solver with the PETSc GMRES strategy.
Figure 3: Time-to-solution versus number of cores for the code_aster turbine blade benchmark on Oracle Cloud Infrastructure, comparing the MUMPS direct solver with the PETSc GMRES strategy.

Using the direct MUMPS solver as a baseline, the benchmark shows the expected reduction in time-to-solution as the number of cores increases, but with a progressive saturation of the gains at higher core counts. This behavior is typical of strong scaling studies on sparse finite element workloads. Once the local computational workload per core becomes too small, the relative impact of communication, synchronization, memory traffic, and solver overhead increases significantly. In practice, this means that simply allocating more cores does not guarantee better performance.

The PETSc-based strategy, relying on GMRES with a MUMPS-based single-precision preconditioner, delivered significantly better performance for this benchmark on the tested OCI configuration. Across the explored range, the observed runtimes were consistently lower than with the direct MUMPS-only approach, leading to a substantially improved time-to-solution. This result is particularly interesting because it illustrates a broader HPC principle: on modern cloud infrastructures, hybrid or iterative strategies can outperform purely direct approaches when they better exploit the balance between distributed memory resources, communication patterns, and numerical workload.

At the same time, the results also show that the best configuration is not obtained at the largest core count. Beyond a certain point, the gains flatten and may even partially degrade, indicating that the benchmark enters a regime in which additional parallelism no longer compensates for the induced overhead. For engineering teams, this is a key operational lesson: the goal should not be to maximize resource usage, but to identify the range in which the solver operates most efficiently.

Practical lessons for structural mechanics workloads on cloud HPC

Several practical conclusions can be drawn from this benchmark.

First, for structural mechanics workloads, solver selection is a first-order performance parameter. In non-linear finite element simulations, the cost of repeatedly solving large sparse linear systems often dominates the total runtime. As a result, the choice between a direct approach and a PETSc-based iterative strategy can have a major impact on the overall efficiency of the simulation.

Second, strong scaling behavior must always be interpreted in terms of operational efficiency rather than raw parallelism. Beyond a certain number of cores, additional resources bring limited benefit and may even lead to suboptimal cost-performance ratios. This is especially important in a cloud environment, where each additional compute resource directly impacts the total simulation cost.

Third, these results confirm that benchmarking remains essential before moving to production usage. There is no universal "best" configuration that applies to all structural mechanics workloads. The optimal choice depends on the type of non-linearity, the mesh size, the conditioning of the algebraic system, and the solver strategy. Cloud HPC therefore provides flexibility, but extracting real value from it requires a clear understanding of the numerical behavior of the simulation code.

From CFD to structural mechanics: a common HPC message

Although CFD and structural mechanics rely on different mathematical models and numerical kernels, the benchmark results obtained with code_saturne and code_aster converge toward the same conclusion. Efficient use of cloud HPC infrastructure is not achieved by hardware alone. It requires a careful alignment between the physical problem, the numerical method, the solver configuration, and the underlying parallel architecture.

For engineering teams, this means that moving to the cloud is not simply a matter of infrastructure migration. It is an opportunity to rethink simulation workflows in terms of scalability, robustness, and cost-efficiency. In that respect, EDF's open-source ecosystem provides a strong foundation: mature industrial solvers, robust numerical capabilities, and the flexibility needed to adapt HPC strategies to real engineering use cases.

Vous pourriez également aimer…