options

vmc.mov1 - 2023-11-14 15:06:12 - MAQAO 2.18.0

Help is available by moving the cursor above any symbol or by checking MAQAO website.

Global Metrics

Total Time (s)72.61
Profiled Time (s)71.18
Time in analyzed loops (%)12.3
Time in analyzed innermost loops (%)11.9
Time in user code (%)12.8
Compilation Options Score (%)100
Array Access Efficiency (%)91.5
Potential Speedups
Perfect Flow Complexity1.00
Perfect OpenMP + MPI + Pthread1.00
Perfect OpenMP + MPI + Pthread + Perfect Load Distribution19.3
No Scalar IntegerPotential Speedup1.01
Nb Loops to get 80%6
FP VectorisedPotential Speedup1.03
Nb Loops to get 80%9
Fully VectorisedPotential Speedup1.11
Nb Loops to get 80%11
FP Arithmetic OnlyPotential Speedup1.03
Nb Loops to get 80%6

CQA Potential Speedups Summary

Loop Based Profile⏎

Innermost Loop Based Profile⏎

Application Categorization⏎

Compilation Options⏎

Source ObjectIssue
▼vmc.mov1–
○jastrow4e.f
○optci.f
○multideterminante.f
○determinant.f
○optorb.f
○optjas.f
○determinant_psit.f
○determinante.f
○scale_dist.f
○orbitals.f
○multiply_slmi_mderiv.f
○determinante_psit.f
○nonlpsi.f
○deriv_nonlpsi.f
○metrop_mov1_slat.f
○basis_fns.f
○deriv_jastrow4.f90
○optwf_sr.f90
○set_input_data.f90
○splfit.f
○deriv_nonloc.f
○get_norbterm.f90
○detsav.f
○distances.f
○nonloc.f
○slm.f90
○multideterminant.f

Loop Path Count Profile⏎

Cumulated Speedup If No Scalar Integer⏎

Cumulated Speedup If FP Vectorized⏎

Cumulated Speedup If Fully Vectorized⏎

Cumulated Speedup If FP Arithmetic Only⏎

Experiment Summary

Application/home/kcamus/trex/champ/champ/bin/vmc.mov1
Timestamp2023-11-14 15:06:12 Universal Timestamp1699974372
Number of processes observed1 Number of threads observed129
Experiment TypeOpenMP;
Machineip-172-31-68-94
Model NameAMD EPYC 9R14 96-Core Processor
Architecturex86_64 Micro ArchitectureZEN_V4
Cache Size1024 KB Number of Cores96
OS VersionLinux 6.2.0-1015-aws #15~22.04.1-Ubuntu SMP Fri Oct 6 21:37:24 UTC 2023
Architecture used during static analysisx86_64 Micro Architecture used during static analysisZEN_V4
Frequency Driveracpi-cpufreq Frequency Governorperformance
Huge Pagesmadvise Hyperthreadingoff
Number of sockets2 Number of cores per socket96
Compilation Optionsvmc.mov1: F90 Flang - 1.5 2017-05-01 '+flang -DTARGET_ARCHITECTURE=\"avx512\" -DVECTORIZATION=\"avx512\" -I/home/kcamus/trex/champ/champ/buildflang/src/module -I/home/kcamus/trex/champ/champ/buildflang/src/parser -march=native -O2 -cpp -mcmodel=large -ffree-line-length-none -g -fno-omit-frame-pointer -fPIC -D_MPI_ -DCLUSTER -ffixed-form -ffixed-line-length-132 -c -o -I/home/kcamus/openmpi/openmpi-5.0.0/_install/include -I/home/kcamus/openmpi/openmpi-5.0.0/_install/lib'

Configuration Summary

Dataset
Run Command<executable> -i vmc_optimization_15000.inp
Number Processes1
Number Nodes1
Filter{type = number ; value = 10 ; }
Profile Start{unit = none ; value = 0 ; }
×