Help is available by moving the cursor above any
symbol or by checking MAQAO website.
- r0: 8x1
- r1: 8x2
- r2: 8x4
- r3: 8x8
- r4: 8x16
- r5: 8x24
| Metric | r0 | r1 | r2 | r3 | r4 | r5 |
|---|
| Total Time (s) | 282.02 | 142.02 | 74.64 | 40.08 | 25.35 | 20.99 |
| Max (Thread Active Time) (s) | 275.98 | 138.99 | 72.20 | 37.95 | 23.26 | 19.00 |
| Average Active Time (s) | 275.63 | 138.72 | 72.02 | 37.76 | 23.19 | 18.95 |
| Activity Ratio (%) | 100 | 99.9 | 100.0 | 99.9 | 100 | 100 |
| Average number of active threads | 7.990 | 15.968 | 31.924 | 63.686 | 127.600 | 191.521 |
| Affinity Stability (%) | 100 | 100 | 100 | 100 | 100 | 100 |
| GFLOPS | 44.792 | 88.919 | 169.236 | 315.052 | 498.018 | 601.454 |
| Time in analyzed loops (%) | 98.7 | 98.4 | 95.5 | 94.3 | 92.7 | 92.4 |
| Time in analyzed innermost loops (%) | 27.8 | 27.8 | 26.9 | 26.9 | 27.4 | 26.6 |
| Time in user code (%) | 98.8 | 98.5 | 95.6 | 94.3 | 92.7 | 92.4 |
| Compilation Options Score (%) | 66.7 | 66.7 | 66.7 | 66.7 | 66.7 | 66.7 |
| Array Access Efficiency (%) | 59.2 | 59.2 | 59.2 | 59.5 | 60.2 | 59.7 |
|
| Potential Speedups |  |
| Perfect Flow Complexity | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| Perfect OpenMP + MPI + Pthread | 1.01 | 1.01 | 1.02 | 1.03 | 1.04 | 1.04 |
| Perfect OpenMP + MPI + Pthread + Perfect Load Distribution | 1.01 | 1.02 | 1.05 | 1.06 | 1.08 | 1.08 |
| Scalability - Gap | 1.00 | 1.01 | 1.06 | 1.14 | 1.44 | 1.79 |
| No Scalar Integer | Potential Speedup | 1.48 | 1.47 | 1.45 | 1.44 | 1.42 | 1.43 |
| Nb Loops to get 80% | 1 | 1 | 1 | 1 | 1 | 1 |
| FP Vectorised | Potential Speedup | 2.14 | 2.13 | 2.05 | 2.01 | 1.95 | 1.98 |
| Nb Loops to get 80% | 2 | 2 | 2 | 2 | 2 | 2 |
| Fully Vectorised | Potential Speedup | 8.33 | 8.17 | 6.72 | 6.24 | 5.70 | 5.66 |
| Nb Loops to get 80% | 4 | 4 | 4 | 4 | 4 | 3 |
| Only FP Arithmetic | Potential Speedup | 1.84 | 1.83 | 1.79 | 1.77 | 1.75 | 1.74 |
| Nb Loops to get 80% | 1 | 1 | 1 | 1 | 1 | 1 |
| OpenMP perfectly balanced | Potential Speedup | 1.00 | 1.00 | 1.00 | 1.01 | 1.02 | 1.02 |
| Nb Loops to get 80% | 1 | 2 | 3 | 3 | 3 | 3 |
| Source Object | Issue |
| ▼kripke_aocc_v2– | |
| ▼ParallelComm.cpp– | |
| ○ | -march=(target) is missing. |
| ▼Collapse.hpp– | |
| ○ | -march=(target) is missing. |
| Source Object | Issue |
| ▼kripke_aocc_v2– | |
| ▼ParallelComm.cpp– | |
| ○ | -march=(target) is missing. |
| ▼Collapse.hpp– | |
| ○ | -march=(target) is missing. |
| Source Object | Issue |
| ▼kripke_aocc_v2– | |
| ▼ParallelComm.cpp– | |
| ○ | -march=(target) is missing. |
| ▼Collapse.hpp– | |
| ○ | -march=(target) is missing. |
| Source Object | Issue |
| ▼kripke_aocc_v2– | |
| ▼ParallelComm.cpp– | |
| ○ | -march=(target) is missing. |
| ▼Collapse.hpp– | |
| ○ | -march=(target) is missing. |
| Source Object | Issue |
| ▼kripke_aocc_v2– | |
| ▼ParallelComm.cpp– | |
| ○ | -march=(target) is missing. |
| ▼Collapse.hpp– | |
| ○ | -march=(target) is missing. |
| Source Object | Issue |
| ▼kripke_aocc_v2– | |
| ▼ParallelComm.cpp– | |
| ○ | -march=(target) is missing. |
| ▼Collapse.hpp– | |
| ○ | -march=(target) is missing. |
| r0 | r1 | r2 | r3 | r4 | r5 |
| Experiment Name | | | | | | |
| Application | /beegfs/hackathon/users/eoseret/kripke_aocc_v2 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Timestamp | 2025-05-14 12:18:56 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Experiment Type | MPI; | MPI; OpenMP; | same as r1 | same as r1 | same as r1 | same as r1 |
| Machine | gmz12.benchmarkcenter.megware.com | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Architecture | x86_64 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Micro Architecture | ZEN_V5 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Model Name | AMD EPYC 9655 96-Core Processor | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Cache Size | 1024 KB | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Number of Cores | 96 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Maximal Frequency | 4.509375 GHz | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| OS Version | Linux 5.14.0-503.31.1.el9_5.x86_64 #1 SMP PREEMPT_DYNAMIC Thu Mar 13 06:50:51 EDT 2025 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Architecture used during static analysis | x86_64 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Micro Architecture used during static analysis | ZEN_V5 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Compilation Options |
kripke_aocc_v2: AMD clang version 17.0.6 (CLANG: AOCC_5.0.0-Build#1377 2024_09_24) /home/eoseret/aocc-compiler-5.0.0/bin/clang-17 --driver-mode=g++ -I /beegfs/hackathon/users/eoseret/Kripke/src -I /beegfs/hackathon/users/eoseret/Kripke/build/include -I /beegfs/hackathon/users/eoseret/Kripke/tpl/raja/include -I /beegfs/hackathon/users/eoseret/Kripke/build/tpl/raja/include -I /beegfs/hackathon/users/eoseret/Kripke/tpl/raja/tpl/camp/include -I /beegfs/hackathon/users/eoseret/Kripke/build/tpl/raja/tpl/camp/include -isystem /cluster/intel/oneapi/2024.0.0/mpi/2021.11/include -g -grecord-command-line -fno-omit-frame-pointer -O3 -D NDEBUG -std=c++14 -fPIC -fopenmp=libomp -MD -MT CMakeFiles/kripke.dir/src/Kripke/Kernel/SweepSubdomain.cpp.o -MF CMakeFiles/kripke.dir/src/Kripke/Kernel/SweepSubdomain.cpp.o.d -o CMakeFiles/kripke.dir/src/Kripke/Kernel/SweepSubdomain.cpp.o -c /beegfs/hackathon/users/eoseret/Kripke/src/Kripke/Kernel/SweepSubdomain.cpp | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Number of processes observed | 8 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Number of threads observed | 8 | 16 | 32 | 64 | 128 | 192 |
| Frequency Driver | acpi-cpufreq | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Frequency Governor | ondemand | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Huge Pages | always | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Hyperthreading | on | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Number of sockets | 2 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Number of cores per socket | 96 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| MAQAO version | 2.21.4 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| MAQAO build | 9c55c8b14007b328890413aa6ef8916a305281ee::20250512-200501 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Comments | | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |