Help is available by moving the cursor above any
symbol or by checking MAQAO website.
| Metric | r0 | r1 | r2 | r3 | r4 | r5 | |
|---|---|---|---|---|---|---|---|
| Total Time (s) | 50.51 | 62.85 | 68.07 | 68.10 | 68.85 | 69.24 | |
| Max (Thread Active Time) (s) | 42.25 | 54.80 | 60.00 | 59.95 | 60.69 | 61.13 | |
| Average Active Time (s) | 42.20 | 54.37 | 59.52 | 59.45 | 60.15 | 60.57 | |
| Activity Ratio (%) | 97.4 | 98.9 | 99.0 | 99.0 | 99.0 | 98.9 | |
| Average number of active threads | 6.684 | 55.360 | 83.940 | 111.754 | 139.796 | 167.967 | |
| Affinity Stability (%) | 97.3 | 99.5 | 99.6 | 99.7 | 99.7 | 99.5 | |
| GFLOPS | 50.926 | 39.586 | 36.189 | 36.245 | 35.831 | 35.604 | |
| Time in analyzed loops (%) | 90.0 | 62.1 | 62.9 | 61.7 | 61.5 | 61.9 | |
| Time in analyzed innermost loops (%) | 88.9 | 62.0 | 62.7 | 61.6 | 61.4 | 61.8 | |
| Time in user code (%) | 91.5 | 62.4 | 63.1 | 61.9 | 61.8 | 62.2 | |
| Compilation Options Score (%) | 99.6 | 99.8 | 99.8 | 99.8 | 99.8 | 99.8 | |
| Array Access Efficiency (%) | 100 | 100 | 100 | 100 | 100 | 100 | |
| Potential Speedups | |||||||
| Perfect Flow Complexity | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | |
| Perfect OpenMP/MPI/Pthread/TBB | 1.05 | 1.05 | 1.04 | 1.03 | 1.05 | 1.04 | |
| Perfect OpenMP/MPI/Pthread/TBB + Perfect Load Distribution | 1.09 | 1.61 | 1.59 | 1.62 | 1.63 | 1.62 | |
| Scalability - Gap | 1.00 | 9.95 | 16.17 | 21.57 | 27.26 | 32.90 | |
| No Scalar Integer | Potential Speedup | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| Nb Loops to get 80% | 1 | 1 | 1 | 1 | 1 | 1 | |
| FP Vectorised | Potential Speedup | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| Nb Loops to get 80% | 1 | 1 | 1 | 1 | 1 | 1 | |
| Fully Vectorised | Potential Speedup | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| Nb Loops to get 80% | 3 | 3 | 2 | 2 | 1 | 1 | |
| Only FP Arithmetic | Potential Speedup | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| Nb Loops to get 80% | 2 | 2 | 2 | 2 | 2 | 2 | |
| OpenMP perfectly balanced | Potential Speedup | 1.08 | 1.61 | 1.58 | 1.61 | 1.61 | 1.60 |
| Nb Loops to get 80% | 1 | 1 | 1 | 1 | 1 | 1 | |
| Source Object | Issue |
|---|---|
| ▼libllama.so | |
| ▼ | |
| ○ | -g is missing for some functions (possibly ones added by the compiler), it is needed to have more accurate reports. Other recommended flags are: -O2/-O3, -march=(target) |
| ○ | -O2, -O3 or -Ofast is missing. |
| ○ | -march=(target) is missing. |
| ▼exec | |
| ▼ | |
| ○ | -g is missing for some functions (possibly ones added by the compiler), it is needed to have more accurate reports. Other recommended flags are: -O2/-O3, -march=(target) |
| ○ | -O2, -O3 or -Ofast is missing. |
| ○ | -march=(target) is missing. |
| ▼libggml-base.so | |
| ▼ | |
| ○ | -g is missing for some functions (possibly ones added by the compiler), it is needed to have more accurate reports. Other recommended flags are: -O2/-O3, -march=(target) |
| ○ | -O2, -O3 or -Ofast is missing. |
| ○ | -march=(target) is missing. |
| ▼libggml-cpu.so | |
| ▼binary-ops.cpp | |
| ○ | |
| ▼ops.cpp | |
| ○ | |
| ▼sgemm.cpp | |
| ○ | |
| ▼vec.cpp | |
| ○ | |
| ▼ggml-cpu.c | |
| ○ | |
| ▼quants.c | |
| ○ |
| r0 | r1 | r2 | r3 | r4 | r5 | |
|---|---|---|---|---|---|---|
| Application | /beegfs/hackathon/users/eoseret/qaas_runs_test/175-950-2189/intel/llama.cpp/run/binaries/icx_3/exec | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Timestamp | 2025-10-06 13:14:44 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Experiment Type | MPI; OpenMP; | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Machine | gmz12.benchmarkcenter.megware.com | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Architecture | x86_64 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Micro Architecture | ZEN_V5 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Model Name | AMD EPYC 9655 96-Core Processor | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Cache Size | 1024 KB | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Number of Cores | 96 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Maximal Frequency | 4.509375 GHz | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| OS Version | Linux 5.14.0-570.39.1.el9_6.x86_64 #1 SMP PREEMPT_DYNAMIC Thu Sep 4 05:08:52 EDT 2025 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Architecture used during static analysis | x86_64 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Micro Architecture used during static analysis | ZEN_V5 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Compilation Options | exec: N/A libggml-base.so: N/A libggml-cpu.so: clang based Intel(R) oneAPI DPC++/C++ Compiler 2025.1.0 (2025.1.0.20250317) /cluster/intel/oneapi/2025.1.0/compiler/2025.1/bin/compiler/clang --intel -D GGML_BACKEND_BUILD -D GGML_BACKEND_SHARED -D GGML_SCHED_MAX_COPIES=4 -D GGML_SHARED -D GGML_USE_CPU_REPACK -D GGML_USE_LLAMAFILE -D GGML_USE_OPENMP -D _GNU_SOURCE -D _XOPEN_SOURCE=600 -D ggml_cpu_EXPORTS -I /beegfs/hackathon/users/eoseret/qaas_runs_test/175-950-2189/intel/llama.cpp/build/llama.cpp/ggml/src/.. -I /beegfs/hackathon/users/eoseret/qaas_runs_test/175-950-2189/intel/llama.cpp/build/llama.cpp/ggml/src/. -I /beegfs/hackathon/users/eoseret/qaas_runs_test/175-950-2189/intel/llama.cpp/build/llama.cpp/ggml/src/ggml-cpu -I /beegfs/hackathon/users/eoseret/qaas_runs_test/175-950-2189/intel/llama.cpp/build/llama.cpp/ggml/src/../include -O3 -O3 -march=znver5 -axCORE-AVX512 -mprefer-vector-width=256 -g -fno-omit-frame-pointer -fcf-protection=none -no-pie -grecord-command-line -fno-finite-math-only -O3 -D NDEBUG -std=gnu11 -fPIC -Wshadow -Wstrict-prototypes -Wpointer-arith -Wmissing-prototypes -Werror=implicit-int -Werror=implicit-function-declaration -Wall -Wextra -Wpedantic -Wcast-qual -Wno-unused-function -fno-associative-math -fiopenmp -MD -MT ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/arch/x86/quants.c.o -MF ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/arch/x86/quants.c.o.d -o ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/arch/x86/quants.c.o -c /beegfs/hackathon/users/eoseret/qaas_runs_test/175-950-2189/intel/llama.cpp/build/llama.cpp/ggml/src/ggml-cpu/arch/x86/quants.c -fveclib=SVML libllama.so: N/A | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Number of processes observed | 1 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Number of threads observed | 8 | 64 | 96 | 128 | 160 | 192 |
| Frequency Driver | acpi-cpufreq | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Frequency Governor | performance | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Huge Pages | always | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Hyperthreading | on | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Number of sockets | 2 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Number of cores per socket | 96 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| MAQAO version | 2025.1.2 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| MAQAO build | ad4b42c12cfbc289a7a711f3ded92abe2eb90c0a::20250917-142411 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |
| Comments | OV scalability run using icx_3 | same as r0 | same as r0 | same as r0 | same as r0 | same as r0 |