options

miniqmc - 2023-06-18 18:22:01 - MAQAO 2.17.4

Help is available by moving the cursor above any symbol or by checking MAQAO website.

Global Metrics

Total Time (s)35.49
Profiled Time (s)34.91
Time in analyzed loops (%)87.2
Time in analyzed innermost loops (%)87.1
Time in user code (%)88.9
Compilation Options Score (%)100
Perfect Flow Complexity1.00
Iterations Count1.16
Array Access Efficiency (%)54.0
Perfect OpenMP + MPI + Pthread1.01
Perfect OpenMP + MPI + Pthread + Perfect Load Distribution1.03
No Scalar IntegerPotential Speedup1.08
Nb Loops to get 80%2
FP VectorisedPotential Speedup1.19
Nb Loops to get 80%2
Fully VectorisedPotential Speedup2.13
Nb Loops to get 80%3
Data In L1 CachePotential Speedup2.35
Nb Loops to get 80%2
FP Arithmetic OnlyPotential Speedup1.60
Nb Loops to get 80%2

CQA Potential Speedups Summary

Loop Based Profile⏎

Innermost Loop Based Profile⏎

Application Categorization⏎

Compilation Options⏎

Source ObjectIssue
▼miniqmc–
○MultiBsplineEvalHelper.hpp
○WaveFunction.cpp
○einspline_spo_omp.cpp
○NewTimer.cpp
○stl_algobase.h
○SoaDistanceTableAAOMPTarget.h
○DiracMatrix.h
○ParticleSet.cpp
○TwoBodyJastrow.h
○OneBodyJastrow.h
○NonLocalPP.hpp
○OhmmsVector.h
○ParticleBConds3DSoa.h
○DiracDeterminant.cpp
○SoaDistanceTableABOMPTarget.h
○BsplineFunctor.h
○random.h
○DelayedUpdate.h
○miniqmc.cpp
○MultiBsplineVGLH_OMPoffload.hpp

Loop Path Count Profile⏎

Loop Iteration Count Profile⏎

Cumulated Speedup If No Scalar Integer⏎

Cumulated Speedup If FP Vectorized⏎

Cumulated Speedup If Fully Vectorized⏎

Cumulated Speedup If Data In L1⏎

Cumulated Speedup If FP Arithmetic Only⏎

Experiment Summary

Application./miniqmc
Timestamp2023-06-18 18:22:01 Universal Timestamp1687105321
Number of processes observed1 Number of threads observed16
Experiment TypeOpenMP;
Machineskylake
Model NameIntel(R) Xeon(R) Platinum 8170 CPU @ 2.10GHz
Architecturex86_64 Micro ArchitectureSKYLAKE
Cache Size36608 KB Number of Cores26
OS VersionLinux 6.2.12-arch1-1 #1 SMP PREEMPT_DYNAMIC Thu, 20 Apr 2023 16:11:55 +0000
Architecture used during static analysisx86_64 Micro Architecture used during static analysisSKYLAKE
Frequency Driverintel_cpufreq Frequency Governorschedutil
Huge Pagesalways Hyperthreadingoff
Number of sockets2 Number of cores per socket26
Compilation Options
miniqmc: clang based Intel(R) oneAPI DPC++/C++ Compiler 2023.0.0 (2023.0.0.20221201) --driver-mode=g++ --intel -I /home/eoseret/miniqmc/src -I /home/eoseret/miniqmc/build_icx/src -I /home/eoseret/miniqmc/src/Particle -I /home/eoseret/miniqmc/src/Utilities -I /home/eoseret/miniqmc/src/Platforms -I /home/eoseret/miniqmc/src/Platforms/Host -D ADD_ -D H5_USE_16_API -D HAVE_CONFIG_H -D HAVE_MKL -D OPENMP_NO_COMPLEX -D restrict=__restrict__ -isystem /opt/intel/oneapi/mkl/2023.0.0/include -fno-omit-frame-pointer -qopt-zmm-usage=high -fiopenmp -fstrict-aliasing -march=native -O2 -g -D NDEBUG -std=c++17 -MD -MT src/QMCWaveFunctions/CMakeFiles/qmcwfs.dir/einspline_spo_omp.cpp.o -MF CMakeFiles/qmcwfs.dir/einspline_spo_omp.cpp.o.d -o CMakeFiles/qmcwfs.dir/einspline_spo_omp.cpp.o -c /home/eoseret/miniqmc/src/QMCWaveFunctions/einspline_spo_omp.cpp -fveclib=SVML -fheinous-gnu-extensions
Commentsminiqmc compiled with icx -qopt-zmm-usage=high, run on Skylake-SP using 16 threads

Configuration Summary

Dataset
Run Command<executable> -g "2 2 2"
Number Processes1
Number Nodes1
Filter{type = number ; value = 10 ; }
Profile StartNot Used
Maximal Path Number4
×