| Message ID | 20260813094535.10083-3-gautam@linux.ibm.com (mailing list archive) |
|---|---|
| State | New |
| Headers | show |
| Series | powerpc/perf: Fix vpa-pmu driver to report correct numbers | expand |
On Thu, Aug 13, 2026 at 09:58:50AM +0000, sashiko-bot@kernel.org wrote: > Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider: > > New issues: > - [High] Task-bound PMU events cache a stale vCPU pointer, resulting in tracking the wrong vCPU. > - [High] CPU-wide and system-wide PMU events are broken and silently report 0 if added when no guest is running. > - [High] Uninitialized event state on task migration leaks stale vCPU pointers. > - [Medium] Unbalanced disablement of the global hardware PMU counters. > > Pre-existing issues: > - [Critical] Dereferencing `kvm_hstate.kvm_vcpu` leads to a Use-After-Free if the previously running VM has been destroyed. > -- > > --- Patch [2]: [PATCH v2 2/3] powerpc/perf: Use the aggregate context switch values from vcpu struct --- > commit e328e1f6c8baf0140410c1983f559e87d1d0efdc > Author: Gautam Menghani <gautam@linux.ibm.com> > powerpc/perf: Use the aggregate context switch values from vcpu struct > > This commit aims to fix inaccurate vpa-pmu counter reporting during vCPU > migrations and when vCPUs are not running by fetching aggregate values > from the vcpu struct instead of relying on the host cpu's PACA. > > > diff --git a/arch/powerpc/perf/vpa-pmu.c b/arch/powerpc/perf/vpa-pmu.c > > --- a/arch/powerpc/perf/vpa-pmu.c > > +++ b/arch/powerpc/perf/vpa-pmu.c > [ ... ] > > @@ -99,19 +99,19 @@ static unsigned long get_counter_data(struct perf_event *event) > > switch (config) { > > case L1_TO_L2_CS_LAT: > > if (event->attach_state & PERF_ATTACH_TASK) > > - data = kvmhv_get_l1_to_l2_cs_time_vcpu(); > > + data = vcpu->arch.l1_to_l2_cs; > > [Severity: Critical] > This is a pre-existing issue, but does dereferencing the vcpu pointer here > lead to a use-after-free if the previously running VM has been destroyed? > > When KVM exits a guest, it does not clear local_paca->kvm_hstate.kvm_vcpu, > leaving a dangling pointer in the host's PACA. > > The PMU driver reads this dangling pointer in vpa_pmu_add() and permanently > caches it in event->pmu_private. Subsequent calls to vpa_pmu_read() pass > this cached pointer to get_counter_data(), which dereferences the freed > memory here. > > > else > > data = kvmhv_get_l1_to_l2_cs_time(); > > break; > [ ... ] > > @@ -126,21 +126,33 @@ static int vpa_pmu_add(struct perf_event *event, int flags) > > { > > u64 data; > > + struct kvm_vcpu *vcpu; > > > > + vcpu = local_paca->kvm_hstate.kvm_vcpu; > > + if (!vcpu) > > + goto out; > A dangling pointer is stored, but will that cause a UAF? When tracing a particular tid/pid, the perf stat command exits if the pid (guest) is destroyed. > [Severity: High] > Does this early exit break CPU-wide and system-wide profiling? > > If a user initiates a CPU-wide profiling session (e.g. perf stat -a) and no > guest is actively running on that specific CPU at that moment, vcpu is NULL. > The function exits early returning 0 without initializing event->pmu_private > or event->hw.prev_count. > > Subsequent calls to vpa_pmu_read() check if (!vcpu) and immediately return, > silently leaving the event counter permanently at 0. The vpa-pmu driver is only meant to report context switch times of individual vCPUs. So reporting 0 when a vCPU is not running is expected behaviour. > > [Severity: High] > Could this early exit also leak state from a previous CPU when a task > migrates? > > When a task with a vpa_pmu event migrates to a new CPU, perf core calls > pmu->del on the old CPU and pmu->add on the new CPU. If vcpu is NULL on the > new CPU, it returns 0 (success) but jumps to out without updating or > clearing event->pmu_private. > > The event silently retains the old vcpu pointer and prev_count from the > previous CPU, resulting in corrupted profiling data calculated against an > unrelated CPU's context. Yes, valid point. > > [Severity: Medium] > Are the global hardware PMU counters skipped here on early bailout, leading > to an unbalanced disablement later? > > If a user creates a PMU event on a CPU where no vCPU is active, this early > exit skips calling kvmhv_set_l2_counters_status(cpu, true). > > However, when the event is deleted or scheduled out, vpa_pmu_del() > unconditionally calls kvmhv_set_l2_counters_status(cpu, false). This > unbalanced behavior forcefully disables the hardware counters even when they > were never enabled, which could break legitimate concurrent KVM tracing > sessions. Valid find - in vpa_pmu_del() the counters should be disabled only if they were enabled in vpa_pmu_add() > > > + > > + event->pmu_private = vcpu; > > [Severity: High] > Is it possible caching the vcpu pointer here results in tracking the wrong > vCPU for task-bound events? > > When a user creates a task-bound vpa_pmu perf event, the core perf > subsystem calls vpa_pmu_add() during perf_event_context_sched_in(). At this > point, the task has been scheduled but has not yet invoked the KVM_RUN > ioctl. > > Because of this, local_paca->kvm_hstate.kvm_vcpu does not point to the > incoming task's vCPU; it contains a stale pointer to the last vCPU that > executed on this physical host CPU. Caching this pointer permanently means > the event tracks unrelated data. Valid find, will fix > > > kvmhv_set_l2_counters_status(smp_processor_id(), true); > > > > - data = get_counter_data(event); > > + data = get_counter_data(vcpu, event); > > local64_set(&event->hw.prev_count, data); > > > > +out: > > return 0; > > } > > -- > Sashiko AI review ยท https://sashiko.dev/#/patchset/20260813094535.10083-1-gautam@linux.ibm.com?part=2
diff --git a/arch/powerpc/include/asm/kvm_book3s_64.h b/arch/powerpc/include/asm/kvm_book3s_64.h index b936e174eefd..f620e3126d68 100644 --- a/arch/powerpc/include/asm/kvm_book3s_64.h +++ b/arch/powerpc/include/asm/kvm_book3s_64.h @@ -691,9 +691,6 @@ void kvmhv_set_l2_counters_status(int cpu, bool status); u64 kvmhv_get_l1_to_l2_cs_time(void); u64 kvmhv_get_l2_to_l1_cs_time(void); u64 kvmhv_get_l2_runtime_agg(void); -u64 kvmhv_get_l1_to_l2_cs_time_vcpu(void); -u64 kvmhv_get_l2_to_l1_cs_time_vcpu(void); -u64 kvmhv_get_l2_runtime_agg_vcpu(void); #endif /* CONFIG_KVM_BOOK3S_HV_POSSIBLE */ diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c index 342168b8bfc8..b9285b7f1fed 100644 --- a/arch/powerpc/kvm/book3s_hv.c +++ b/arch/powerpc/kvm/book3s_hv.c @@ -4189,51 +4189,6 @@ u64 kvmhv_get_l2_runtime_agg(void) } EXPORT_SYMBOL(kvmhv_get_l2_runtime_agg); -u64 kvmhv_get_l1_to_l2_cs_time_vcpu(void) -{ - struct kvm_vcpu *vcpu; - struct kvm_vcpu_arch *arch; - - vcpu = local_paca->kvm_hstate.kvm_vcpu; - if (vcpu) { - arch = &vcpu->arch; - return arch->l1_to_l2_cs; - } else { - return 0; - } -} -EXPORT_SYMBOL(kvmhv_get_l1_to_l2_cs_time_vcpu); - -u64 kvmhv_get_l2_to_l1_cs_time_vcpu(void) -{ - struct kvm_vcpu *vcpu; - struct kvm_vcpu_arch *arch; - - vcpu = local_paca->kvm_hstate.kvm_vcpu; - if (vcpu) { - arch = &vcpu->arch; - return arch->l2_to_l1_cs; - } else { - return 0; - } -} -EXPORT_SYMBOL(kvmhv_get_l2_to_l1_cs_time_vcpu); - -u64 kvmhv_get_l2_runtime_agg_vcpu(void) -{ - struct kvm_vcpu *vcpu; - struct kvm_vcpu_arch *arch; - - vcpu = local_paca->kvm_hstate.kvm_vcpu; - if (vcpu) { - arch = &vcpu->arch; - return arch->l2_runtime_agg; - } else { - return 0; - } -} -EXPORT_SYMBOL(kvmhv_get_l2_runtime_agg_vcpu); - #else int kvmhv_get_l2_counters_status(void) { diff --git a/arch/powerpc/perf/vpa-pmu.c b/arch/powerpc/perf/vpa-pmu.c index bff4cfab7b94..e79d98447c74 100644 --- a/arch/powerpc/perf/vpa-pmu.c +++ b/arch/powerpc/perf/vpa-pmu.c @@ -91,7 +91,7 @@ static int vpa_pmu_event_init(struct perf_event *event) return 0; } -static unsigned long get_counter_data(struct perf_event *event) +static unsigned long get_counter_data(struct kvm_vcpu *vcpu, struct perf_event *event) { unsigned int config = event->attr.config; u64 data; @@ -99,19 +99,19 @@ static unsigned long get_counter_data(struct perf_event *event) switch (config) { case L1_TO_L2_CS_LAT: if (event->attach_state & PERF_ATTACH_TASK) - data = kvmhv_get_l1_to_l2_cs_time_vcpu(); + data = vcpu->arch.l1_to_l2_cs; else data = kvmhv_get_l1_to_l2_cs_time(); break; case L2_TO_L1_CS_LAT: if (event->attach_state & PERF_ATTACH_TASK) - data = kvmhv_get_l2_to_l1_cs_time_vcpu(); + data = vcpu->arch.l2_to_l1_cs; else data = kvmhv_get_l2_to_l1_cs_time(); break; case L2_RUNTIME_AGG: if (event->attach_state & PERF_ATTACH_TASK) - data = kvmhv_get_l2_runtime_agg_vcpu(); + data = vcpu->arch.l2_runtime_agg; else data = kvmhv_get_l2_runtime_agg(); break; @@ -126,21 +126,33 @@ static unsigned long get_counter_data(struct perf_event *event) static int vpa_pmu_add(struct perf_event *event, int flags) { u64 data; + struct kvm_vcpu *vcpu; + vcpu = local_paca->kvm_hstate.kvm_vcpu; + if (!vcpu) + goto out; + + event->pmu_private = vcpu; kvmhv_set_l2_counters_status(smp_processor_id(), true); - data = get_counter_data(event); + data = get_counter_data(vcpu, event); local64_set(&event->hw.prev_count, data); +out: return 0; } static void vpa_pmu_read(struct perf_event *event) { u64 prev_data, new_data, final_data; + struct kvm_vcpu *vcpu; + + vcpu = (struct kvm_vcpu *) event->pmu_private; + if (!vcpu) + return; prev_data = local64_read(&event->hw.prev_count); - new_data = get_counter_data(event); + new_data = get_counter_data(vcpu, event); final_data = new_data - prev_data; local64_add(final_data, &event->count);
The vpa-pmu driver reports incorrect numbers in 2 scenarios: 1. The vCPU process gets rescheduled to a different host cpu - Incorrect numbers are observed here because the PACA is per-host cpu resource, and the KVM vCPUs can be rescheduled to different host cpus. This causes the vpa-pmu driver to subtract wrong values when vCPUs are rescheduled. 2. The vCPU is not running when vpa_pmu_read() is called. - In this case get_counter_data() returns 0, and this can result in negative numbers getting reported. Fix the above issues by using the aggregate values from the vcpu structure to capture and report the difference in counter values. Fixes: 176cda0619b6 ("powerpc/perf: Add perf interface to expose vpa counters") Cc: <stable@vger.kernel.org> Signed-off-by: Gautam Menghani <gautam@linux.ibm.com> --- v1 -> v2: 1. Fix the output for "perf stat --cpu ..." 2. Add stable/fixes tags arch/powerpc/include/asm/kvm_book3s_64.h | 3 -- arch/powerpc/kvm/book3s_hv.c | 45 ------------------------ arch/powerpc/perf/vpa-pmu.c | 24 +++++++++---- 3 files changed, 18 insertions(+), 54 deletions(-)