| Message ID | 20260902062153.31048-3-logang@deltatee.com |
|---|---|
| State | New |
| Headers | show |
| Series | dmaengine: miscellaneous fixes for ioat, switchtec_dma and plx_dma | expand |
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider: Pre-existing issues: - [High] Race condition in switchtec_dma_remove allows concurrent channel allocations on a broken channel. -- commit f369eaa9794c069cc909c5cee24ba55c0cfa8c48 Author: Logan Gunthorpe <logang@deltatee.com> dmaengine: switchtec-dma: fix resource leak in alloc_chan_resources This commit adds an error-unwind path to switchtec_dma_alloc_chan_resources() to prevent resource leaks when channel initialization fails. It ensures descriptor rings are freed and the channel is disabled and marked inactive on error. Link: https://lore.kernel.org/dmaengine/20260707165555.350951F000E9@smtp.kernel.org > diff --git a/drivers/dma/switchtec_dma.c b/drivers/dma/switchtec_dma.c > index a10818efba4e5..e55fe9ff7e2c5 100644 > --- a/drivers/dma/switchtec_dma.c > +++ b/drivers/dma/switchtec_dma.c [ ... ] > @@ -1007,7 +1007,8 @@ static int switchtec_dma_alloc_chan_resources(struct dma_chan *chan) > rcu_read_lock(); > if (!rcu_dereference(swdma_dev->pdev)) { [Severity: High] This isn't a bug introduced by this patch, but is there a pre-existing race condition here between channel allocation and device removal? During device unbind or hot-unplug, switchtec_dma_remove() calls switchtec_dma_chans_release(), which frees the channel IRQs and kills tasklets, then clears swdma_dev->pdev: drivers/dma/switchtec_dma.c:switchtec_dma_remove() { ... switchtec_dma_chans_release(pdev, swdma_dev); rcu_assign_pointer(swdma_dev->pdev, NULL); synchronize_rcu(); pci_free_irq(pdev, swdma_dev->chan_status_irq, swdma_dev); pci_free_irq_vectors(pdev); dma_async_device_unregister(&swdma_dev->dma_dev); } Because dma_async_device_unregister() happens much later, a concurrent client calling dma_request_chan() can still invoke switchtec_dma_alloc_chan_resources(). If this executes before swdma_dev->pdev is set to NULL, it successfully allocates the channel and sets ring_active = true, unaware that the IRQ and tasklet were already destroyed. Could this cause the client to hang waiting for a freed IRQ, or trigger a system crash due to unmapped memory access when switchtec_dma_remove() subsequently unmaps the hardware registers? > rcu_read_unlock(); > - return -ENODEV; > + rc = -ENODEV; > + goto err_ring_inactive; > } [ ... ]
diff --git a/drivers/dma/switchtec_dma.c b/drivers/dma/switchtec_dma.c index a10818efba4e..e55fe9ff7e2c 100644 --- a/drivers/dma/switchtec_dma.c +++ b/drivers/dma/switchtec_dma.c @@ -988,15 +988,15 @@ static int switchtec_dma_alloc_chan_resources(struct dma_chan *chan) rc = enable_channel(swdma_chan); if (rc) - return rc; + goto err_free_desc; rc = reset_channel(swdma_chan); if (rc) - return rc; + goto err_disable_channel; rc = unhalt_channel(swdma_chan); if (rc) - return rc; + goto err_disable_channel; swdma_chan->ring_active = true; swdma_chan->comp_ring_active = true; @@ -1007,7 +1007,8 @@ static int switchtec_dma_alloc_chan_resources(struct dma_chan *chan) rcu_read_lock(); if (!rcu_dereference(swdma_dev->pdev)) { rcu_read_unlock(); - return -ENODEV; + rc = -ENODEV; + goto err_ring_inactive; } perf_cfg = readl(&swdma_chan->mmio_chan_fw->perf_cfg); @@ -1029,6 +1030,20 @@ static int switchtec_dma_alloc_chan_resources(struct dma_chan *chan) FIELD_GET(PERF_MRRS_MASK, perf_cfg)); return SWITCHTEC_DMA_SQ_SIZE; + +err_ring_inactive: + spin_lock_bh(&swdma_chan->submit_lock); + swdma_chan->ring_active = false; + spin_unlock_bh(&swdma_chan->submit_lock); + + spin_lock_bh(&swdma_chan->complete_lock); + swdma_chan->comp_ring_active = false; + spin_unlock_bh(&swdma_chan->complete_lock); +err_disable_channel: + disable_channel(swdma_chan); +err_free_desc: + switchtec_dma_free_desc(swdma_chan); + return rc; } static void switchtec_dma_free_chan_resources(struct dma_chan *chan)
switchtec_dma_alloc_chan_resources() returns directly on any later failure, without ever freeing the descriptor rings and coherent DMA memory it just allocated. The dmaengine core does not call device_free_chan_resources() when device_alloc_chan_resources() fails, so the driver has to unwind its own partial state. The device-removed check also runs after ring_active and comp_ring_active have already been set true, so a failure there left the channel marked active despite alloc_chan_resources() reporting failure. Add an error-unwind path that disables the channel and frees the descriptor rings on every failure after allocation. ring_active and comp_ring_active are cleared under the same locks switchtec_dma_free_chan_resources() already uses, since the completion tasklet checks comp_ring_active under complete_lock before touching the completion ring, and a stale IRQ can still be in flight when this unwind path runs. Reported-by: Sashiko <sashiko-bot@kernel.org> Link: https://lore.kernel.org/dmaengine/20260707165555.350951F000E9@smtp.kernel.org Fixes: 30eba9df76ad ("dmaengine: switchtec-dma: Implement hardware initialization and cleanup") Reviewed-by: Frank Li <Frank.Li@nxp.com> Signed-off-by: Logan Gunthorpe <logang@deltatee.com> --- drivers/dma/switchtec_dma.c | 23 +++++++++++++++++++---- 1 file changed, 19 insertions(+), 4 deletions(-)