Unverified Commit 4400dcfd authored by jmusser304's avatar jmusser304 Committed by GitHub
Browse files

Merge pull request #15 from hengjiew/master

Document options for "Greedy" load balancing and reducing ghost particles
parents 7fdc3d4b e0bb1b68
Loading
Loading
Loading
Loading
+1 −1
Original line number Diff line number Diff line
@@ -4,7 +4,7 @@
.. role:: fortran(code)
   :language: fortran

.. _ss:dual_grid:
.. _sec:dual_grid:

Dual Grid Approach
------------------
+4 −0
Original line number Diff line number Diff line
@@ -28,3 +28,7 @@ Options supported by AMReX include:

- Round-robin: sort grids and assign them to ranks in round-robin fashion -- specifically
  FAB ``i`` is owned by CPU ``i % N`` where N is the total number of MPI ranks.

These methods work for both fluid and particle grids if dual-grid is enabled. MFiX-Exa also supports
a Greedy load balancing algorithm for particle grids. It balances the particle counts per rank and
aligns particle grids with fluid grids to minimize the data-transfer between two grids.
 No newline at end of file
+12 −7
Original line number Diff line number Diff line
.. role:: cpp(code)
   :language: c++

.. _Chap:InputsLoadBalancing:

Gridding and Load Balancing
@@ -35,21 +38,23 @@ The following inputs must be preceded by "fabarray_mfiter" and determine how we

The following inputs must be preceded by "particles"

+-------------------+-----------------------------------------------------------------------+-------------+--------------+
+----------------------+-----------------------------------------------------------------------+-------------+--------------+
|                      | Description                                                           |   Type      | Default      |
+===================+=======================================================================+=============+==============+
+======================+=======================================================================+=============+==============+
| max_grid_size_x      | Maximum number of cells at level 0 in each grid in x-direction        |    Int      | 32           |
|                      | for grids in the ParticleBoxArray if dual_grid is true                |             |              |
+-------------------+-----------------------------------------------------------------------+-------------+--------------+
+----------------------+-----------------------------------------------------------------------+-------------+--------------+
| max_grid_size_y      | Maximum number of cells at level 0 in each grid in y-direction        |    Int      | 32           |
|                      | for grids in the ParticleBoxArray if dual_grid is true                |             |              |
+-------------------+-----------------------------------------------------------------------+-------------+--------------+
+----------------------+-----------------------------------------------------------------------+-------------+--------------+
| max_grid_size_z      | Maximum number of cells at level 0 in each grid in z-direction        |    Int      | 32           |
|                      | for grids in the ParticleBoxArray if dual_grid is true.               |             |              |
+-------------------+-----------------------------------------------------------------------+-------------+--------------+
+----------------------+-----------------------------------------------------------------------+-------------+--------------+
| tile_size            | Maximum number of cells in each direction for (logical) tiles         |  IntVect    | 1024000,8,8  |
|                      | in the ParticleBoxArray if dual_grid is true.                         |             |              |
+-------------------+-----------------------------------------------------------------------+-------------+--------------+
+----------------------+-----------------------------------------------------------------------+-------------+--------------+
| reduceGhostParticles | whether to remove unused ghost particles                              |    Bool     | false        |
+----------------------+-----------------------------------------------------------------------+-------------+--------------+

Note that when running a granular simulation, i.e., no fluid phase, :cpp:`mfix.dual_grid` must be 0. Hence,
the :cpp:`particles.max_grid_size` (in each direction) have no meaning. Therefore the fluid grid and tile
@@ -66,7 +71,7 @@ The following inputs must be preceded by "mfix" and determine how we load balanc
| load_balance_fluid   | Only relevant if (dual_grid); if so do we also regrid mesh data       |  Int        | 1            |
+----------------------+-----------------------------------------------------------------------+-------------+--------------+
| load_balance_type    | What strategy to use for load balancing                               |  String     | KnapSack     |
|                      | Options are "KnapSack"or "SFC"                                        |             |              |
|                      | Options are "KnapSack", "SFC", or "Greedy"                            |             |              |
+----------------------+-----------------------------------------------------------------------+-------------+--------------+
| knapsack_weight_type | What weighting function to use if using Knapsack load balancing       |  String     | RunTimeCosts |
|                      | Options are "RunTimeCosts" or "NumParticles""                         |             |              |
+15 −1
Original line number Diff line number Diff line
.. role:: cpp(code)
   :language: c++

Particles on GPUs
==========================

@@ -50,7 +53,18 @@ The final on-grid neighbor list data structure consists of two arrays. First, we

Note that, because of our use of managed memory to store the particle data and the neighbor list, the above code will work when compiled for either CPU or GPU.

The above algorithm deals with constructing a neighbor list for the particles on a single grid. When domain decomposition is used, one must also make copies of particles on adjacent grids, potentially performing the necessary MPI communication for grids associated with other processes. The routines `fillNeighbors`, which computes which particles needed to be ghosted to which grid, and `updateNeighbors`, which copies up-to-date data for particles that have already been ghosted, have also been offloaded to the GPU, using techniques similar to AMReX's `Redistribute` routine. The important thing for users is that calling these functions does not trigger copying data off the GPU.
The above algorithm deals with constructing a neighbor list for the particles on a single grid.
When domain decomposition is used, one must also make copies of particles on adjacent grids,
potentially performing the necessary MPI communication for grids associated with other processes.
The routines `fillNeighbors`, which computes which particles needed to be ghosted to which grid, and `updateNeighbors`,
which copies up-to-date data for particles that have already been ghosted, have also been offloaded to the GPU,
using techniques similar to AMReX's `Redistribute` routine.

Note that when the particles are dense, only a small portion of the ghost particles are neighbors to the particles
inside the grids. AMReX provides a function `selectActualNeighbors` to filter out the ghost particles that will
not be used for building the neighbor list. So the subsequent calls to `updateNeighbors` can avoid transferring
these unused ghost particles, which significantly reduces the communication cost in some tests.
To use this optimization, set `particles.reduceGhostParticles` to :cpp:`true` in MFiX-Exa's inputs.

Once the neighbor list has been constructed, collisions with both particles and walls can easily be processed.