linux

History

Ingo Molnar 3dfabc74c6 perf report: Add per system call overhead histogram Take advantage of call-graph percounter sampling/recording to display a non-trivial histogram: the true, collapsed/summarized cost measurement, on a per system call total overhead basis: aldebaran:~/linux/linux/tools/perf> ./perf record -g -a -f ~/hackbench 10 aldebaran:~/linux/linux/tools/perf> ./perf report -s symbol --syscalls \| head -10 # # (3536 samples) # # Overhead Symbol # ........ ...... # 40.75% [k] sys_write 40.21% [k] sys_read 4.44% [k] do_nmi ... This is done by accounting each (reliable) call-chain that chains back to a given system call to that system call function. [ So in the above example we can see that hackbench spends about 40% of its total time somewhere in sys_write() and 40% somewhere in sys_read(), the rest of the time is spent in user-space. The time is not spent in sys_write() _itself_ but in one of its many child functions. ] Or, a recording of a (source files are already in the page-cache) kernel build: $ perf record -g -m 512 -f -- make -j32 kernel $ perf report -s s --syscalls \| grep '\[k\]' \| grep -v nmi 4.14% [k] do_page_fault 1.20% [k] sys_write 1.10% [k] sys_open 0.63% [k] sys_exit_group 0.48% [k] smp_apic_timer_interrupt 0.37% [k] sys_read 0.37% [k] sys_execve 0.20% [k] sys_mmap 0.18% [k] sys_close 0.14% [k] sys_munmap 0.13% [k] sys_poll 0.09% [k] sys_newstat 0.07% [k] sys_clone 0.06% [k] sys_newfstat 0.05% [k] sys_access 0.05% [k] schedule Shows the true total cost of each syscall variant that gets used during a kernel build. This profile reveals it that pagefaults are the costliest, followed by read()/write(). An interesting detail: timer interrupts cost 0.5% - or 0.5 seconds per 100 seconds of kernel build-time. (this was done with HZ=1000) The summary is done in 'perf report', i.e. in the post-processing stage - so once we have a good call-graph recording, this type of non-trivial high-level analysis becomes possible. Cc: Peter Zijlstra <a.p.zijlstra@chello.nl> Cc: Mike Galbraith <efault@gmx.de> Cc: Paul Mackerras <paulus@samba.org> Cc: Arnaldo Carvalho de Melo <acme@redhat.com> Cc: Linus Torvalds <torvalds@linux-foundation.org> Cc: Frederic Weisbecker <fweisbec@gmail.com> Cc: Pekka Enberg <penberg@cs.helsinki.fi> LKML-Reference: <new-submission> Signed-off-by: Ingo Molnar <mingo@elte.hu>		2009-06-15 15:58:03 +02:00
..
Documentation
util	perf stat: Reorganize output	2009-06-13 13:40:03 +02:00
.gitignore
builtin-annotate.c	perf annotate: Fixes for filename:line displays	2009-06-13 17:51:00 +02:00
builtin-help.c
builtin-list.c
builtin-record.c	perf record: Fix fast task-exit race	2009-06-15 09:08:31 +02:00
builtin-report.c	perf report: Add per system call overhead histogram	2009-06-15 15:58:03 +02:00
builtin-stat.c	perf stat: Enable raw data to be printed	2009-06-13 15:40:35 +02:00
builtin-top.c	perf_counter: Standardize event names	2009-06-11 17:54:15 +02:00
builtin.h
command-list.txt
design.txt	perf_counter: Start documenting HAVE_PERF_COUNTERS requirements	2009-06-12 19:37:30 +02:00
Makefile	perf stat: Enable raw data to be printed	2009-06-13 15:40:35 +02:00
perf.c
perf.h	perf_counter: Add forward/backward attribute ABI compatibility	2009-06-12 14:28:52 +02:00