khkhiu.github.io documentation¶
Add your content using reStructuredText syntax. See the
reStructuredText
documentation for details.
Contents:
- Disabling SMT cut our OpenFOAM HPC Challenge runtime by 17%, on half the cores
- Optimizing OpenFOAM’s occDrivAer Case for the APAC HPC-AI Competition
- 1. What the APAC HPC-AI Competition is
- 2. What OHC-1 is
- 3. The occDrivAer case itself
- 4. Decomposition methods, what they are and how they differ
- 5. Renumbering, what RCM means
- 6. What every top OHC-1 hardware-track submission actually used
- 7. Rank count, the cells/core sweet spot
- 8. Settings to try, summary checklist
- 9. Key takeaways
- References
- occDrivAerStaticMesh on Vanda: A 1-Node Decomposition and MPI-Binding Study
- 0. Background: what problem is being solved, and what the vocabulary means
- 1. Setup
- 2. Decomposition shape sweep
- 3. Discovery: node-to-node contention on a shared cluster
- 4. MPI rank-binding sweep
- 5. A note on OpenFOAM’s built-in profiler
- 6. Summary and recommended configuration
- 7. Recommendations for the scored 4-node run
- Appendix: glossary quick-reference
- Rank placement and GAMG-targeted collective tuning cut occDrivAer steady-state time by 10.9% at 2 nodes
- Test matrix and fixed parameters
- 1. Single-node profiling baseline (context from the prior session)
- 2. Decomposition placement sweep — 4 hierarchical variants (job:
v3, profiled harness) - 3. Decomposition method comparison — scotch vs. hierarchical (job:
v4, profiled harness) - 4. MPI tuning sweep on top of the winning decomposition (job:
v5, profiled harness) - 5. Stacking NUMA pinning + GAMG collectives, and a warm-up transient (job:
stack_final, clean harness — no forced write, no profiling instrumentation) - 6. A confound we checked and ruled out
- 7. Job-to-job hardware/allocation variance
- 8. Scaling efficiency (1 → 2 nodes)
- 9. Resource accounting
- 10. Known issues and implementation notes
- Locked configuration going into 4 nodes