# Large model memory allocation and MPI

**URL:** <https://calculix.discourse.group/t/large-model-memory-allocation-and-mpi/4132>\
**Category:** Analysis issues\
**Created:** [September 28, 2026, 7:45am UTC](https://calculix.discourse.group/t/large-model-memory-allocation-and-mpi/4132 "2026-09-28T07:45:45Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![linth](https://avatars.discourse-cdn.com/v4/letter/l/898d66/32.png) [@linth](https://calculix.discourse.group/u/linth)\
**Post date:** [September 28, 2026, 7:45am UTC](https://calculix.discourse.group/t/large-model-memory-allocation-and-mpi/4132/1 "2026-09-28T07:45:45Z")

</div>

Hi there,

I compiled the 2.23 code with MPI to solve with MKL Pardiso. Running a large problem (\>5M equations) on a cluster environment, I faced the problem that the master node runs out of memory, i.e. I received the error

Using up to 16 cpu(s) for the symmetric stiffness/mass contributions.

\*ERROR in u\_calloc: error allocating memory  
variable=au1, file=mafillsmasmain.c, line=179, num=6264344448, size=8

Looking into mafillsmasmain.c + the manual I found that the size of the variable au1

NNEW(au1,double,(long long)num\_cpus\*(nzs[2]+nzs[1]));

scales with the number of cpus and can be reduced by setting the variable CCX\_NPROC\_STIFFNESS to a lower value. Running the same problem with CCX\_NPROC\_STIFFNESS=4 actually run without memory problems.

As I understand that, the number of equations is thus limited by the memory of the master node for now. Has someone tried to change the code such that the assembly of matrices can be distributed to different nodes with individual memory on a cluster (using MPI)?

Thanks

---

<div class="post-metadata">

**Author:** ![cwoelf](https://avatars.discourse-cdn.com/v4/letter/c/74df32/32.png) [@cwoelf](https://calculix.discourse.group/u/cwoelf)\
**Post date:** [September 29, 2026, 8:50am UTC](https://calculix.discourse.group/t/large-model-memory-allocation-and-mpi/4132/2 "2026-09-29T08:50:49Z")

</div>

Hello linth,

as you have pointed out, the issue here is that CalculiX core only works with shared-memory parallelization, MPI is only concerning some of the solvers.

The `au1` buffer for assembly resides in the memory of the main thread. The idea behind the parallelization of the assembly is that each thread gets its own slice within `au1` to assemble into. When all threads are done, a reduction is performed in serial to get `au`.

To allow for proper distributed-memory parallelization, we would need to employ throughout the whole code data structures for our global matrices which are distributed over a given cluster’s nodes’ memories and also redesign the algorithms to be aware of this. It is a fundamental refactor and not a small undertaking. Currently, it is unfortunately not on the agenda but any existing experiments would of course be interesting.

Best regards

Christoph
