From patchwork Tue Oct 8 08:49:34 2024 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: liuhongt X-Patchwork-Id: 1994101 Return-Path: X-Original-To: incoming@patchwork.ozlabs.org Delivered-To: patchwork-incoming@legolas.ozlabs.org Authentication-Results: legolas.ozlabs.org; dkim=pass (2048-bit key; unprotected) header.d=intel.com header.i=@intel.com header.a=rsa-sha256 header.s=Intel header.b=SI5BQZ9P; dkim-atps=neutral Authentication-Results: legolas.ozlabs.org; spf=pass (sender SPF authorized) smtp.mailfrom=gcc.gnu.org (client-ip=2620:52:3:1:0:246e:9693:128c; helo=server2.sourceware.org; envelope-from=gcc-patches-bounces~incoming=patchwork.ozlabs.org@gcc.gnu.org; receiver=patchwork.ozlabs.org) Received: from server2.sourceware.org (server2.sourceware.org [IPv6:2620:52:3:1:0:246e:9693:128c]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature ECDSA (secp384r1) server-digest SHA384) (No client certificate requested) by legolas.ozlabs.org (Postfix) with ESMTPS id 4XN8rQ4N5pz1xtV for ; Tue, 8 Oct 2024 19:51:18 +1100 (AEDT) Received: from server2.sourceware.org (localhost [IPv6:::1]) by sourceware.org (Postfix) with ESMTP id CA2463858C32 for ; Tue, 8 Oct 2024 08:51:16 +0000 (GMT) X-Original-To: gcc-patches@gcc.gnu.org Delivered-To: gcc-patches@gcc.gnu.org Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.18]) by sourceware.org (Postfix) with ESMTPS id 314823861035 for ; Tue, 8 Oct 2024 08:49:39 +0000 (GMT) DMARC-Filter: OpenDMARC Filter v1.4.2 sourceware.org 314823861035 Authentication-Results: sourceware.org; dmarc=pass (p=none dis=none) header.from=intel.com Authentication-Results: sourceware.org; spf=pass smtp.mailfrom=intel.com ARC-Filter: OpenARC Filter v1.0.0 sourceware.org 314823861035 Authentication-Results: server2.sourceware.org; arc=none smtp.remote-ip=192.198.163.18 ARC-Seal: i=1; a=rsa-sha256; d=sourceware.org; s=key; t=1728377381; cv=none; b=OH5XuXA4pN0OhKuKtUvB6qza4fzDzYO1DgiI802846KtOU2G3KLqgwc+y4I3nhXy/gILyFx+QR1wcJMtZ7XCYMS2Elz+GSSgtKYu5LQR6HMEG1PO1gPoQQtkzxKxqGlASsdIZigI1c7shZItrbULjhG1SZ/FwckA5y2JfpUsgRw= ARC-Message-Signature: i=1; a=rsa-sha256; d=sourceware.org; s=key; t=1728377381; c=relaxed/simple; bh=8nu1GomqbPMbGJ+MHKd2PgnQYqK7dVe7KDCrk6weDS0=; h=DKIM-Signature:From:To:Subject:Date:Message-Id:MIME-Version; b=UCLdXkl8wp0Y0v//I8zyju1NBYthNBhTObqvS40c1FDqUrALjZ2Z39ZrqGZ2ieIA3P+o8SSLLew82g96qd5N4/TYCiwom0t7xeWUg5iH+yC4Yg+l1Xoxzg7+6wJT5B2O/41Jncl44BI/VdX6bAn7sXVR7toAxT2SCYoOAruiQU0= ARC-Authentication-Results: i=1; server2.sourceware.org DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1728377379; x=1759913379; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=8nu1GomqbPMbGJ+MHKd2PgnQYqK7dVe7KDCrk6weDS0=; b=SI5BQZ9PwV4z2e3QtwSVkLp9nnYLPurPWY6N49akYlVS0judi8CtVeyA eIoIejHkgZp+LfNd01gk3YwghLcgPbbty3tImO26nNSYqnWnmZ40bsvk1 uDN8vLTHrHLIElLy6BY5RGZMPpDn9U2okPA3QhZXC7kjlRjkfKbHl7j8B 2M5j1zBGlHybfSb1CguN4iY3Jp3CD0uwkHG+iGINkBD9B+4u3d0m+Uq/n Ooyww+NhBlDoW2pApg+kwa/36XB5Pvd+57WQVo8xzyckfycMWXV0wBYXv mC7Bg/Hvo5edOnkHwbIAMbE7rq57pP9t3b3hXjUEQf4sQ9xwDcMQZPQfl A==; X-CSE-ConnectionGUID: 9x8HDefZRMSFeLV+msjZ9w== X-CSE-MsgGUID: 5twjl8w2S3qzMQIyhPMQyw== X-IronPort-AV: E=McAfee;i="6700,10204,11218"; a="27010906" X-IronPort-AV: E=Sophos;i="6.11,186,1725346800"; d="scan'208";a="27010906" Received: from fmviesa009.fm.intel.com ([10.60.135.149]) by fmvoesa112.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 08 Oct 2024 01:49:38 -0700 X-CSE-ConnectionGUID: 9fnzvK8TQuCpTeK24HgYjw== X-CSE-MsgGUID: wIR1fk69Sp63QjPwiVnGJA== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.11,186,1725346800"; d="scan'208";a="75758627" Received: from shliclel4217.sh.intel.com ([10.239.240.127]) by fmviesa009.fm.intel.com with ESMTP; 08 Oct 2024 01:49:37 -0700 From: liuhongt To: gcc-patches@gcc.gnu.org Cc: crazylht@gmail.com, hjl.tools@gmail.com Subject: [PATCH 1/2] [x86] Add new microarchitecture tune for SRF/GRR/CWF. Date: Tue, 8 Oct 2024 16:49:34 +0800 Message-Id: <20241008084935.3371416-2-hongtao.liu@intel.com> X-Mailer: git-send-email 2.31.1 In-Reply-To: <20241008084935.3371416-1-hongtao.liu@intel.com> References: <20241008084935.3371416-1-hongtao.liu@intel.com> MIME-Version: 1.0 X-Spam-Status: No, score=-12.3 required=5.0 tests=BAYES_00, DKIMWL_WL_HIGH, DKIM_SIGNED, DKIM_VALID, DKIM_VALID_AU, DKIM_VALID_EF, GIT_PATCH_0, KAM_SHORT, SPF_HELO_NONE, SPF_NONE, TXREP autolearn=ham autolearn_force=no version=3.4.6 X-Spam-Checker-Version: SpamAssassin 3.4.6 (2021-04-09) on server2.sourceware.org X-BeenThere: gcc-patches@gcc.gnu.org X-Mailman-Version: 2.1.30 Precedence: list List-Id: Gcc-patches mailing list List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: gcc-patches-bounces~incoming=patchwork.ozlabs.org@gcc.gnu.org For Crestmont, 4-operand vex blendv instructions come from MSROM and is slower than 3-instructions sequence (op1 & mask) | (op2 & ~mask). legacy blendv instruction can still be handled by the decoder. The patch add a new tune which is enabled for all processors except for SRF/CWF. It will use vpand + vpandn + vpor instead of vpblendvb(similar for vblendvps/vblendvpd) for SRF/CWF. gcc/ChangeLog: * config/i386/i386-expand.cc (ix86_expand_sse_movcc): Guard instruction blendv generation under new tune. * config/i386/i386.h (TARGET_SSE_MOVCC_USE_BLENDV): New Macro. * config/i386/x86-tune.def (X86_TUNE_SSE_MOVCC_USE_BLENDV): New tune. --- gcc/config/i386/i386-expand.cc | 24 +++++++++---------- gcc/config/i386/i386.h | 2 ++ gcc/config/i386/x86-tune.def | 8 +++++++ .../gcc.target/i386/sse_movcc_use_blendv.c | 12 ++++++++++ 4 files changed, 34 insertions(+), 12 deletions(-) create mode 100644 gcc/testsuite/gcc.target/i386/sse_movcc_use_blendv.c diff --git a/gcc/config/i386/i386-expand.cc b/gcc/config/i386/i386-expand.cc index 124cb976ec8..e4087cccb7c 100644 --- a/gcc/config/i386/i386-expand.cc +++ b/gcc/config/i386/i386-expand.cc @@ -4254,23 +4254,23 @@ ix86_expand_sse_movcc (rtx dest, rtx cmp, rtx op_true, rtx op_false) switch (mode) { case E_V2SFmode: - if (TARGET_SSE4_1) + if (TARGET_SSE_MOVCC_USE_BLENDV && TARGET_SSE4_1) gen = gen_mmx_blendvps; break; case E_V4SFmode: - if (TARGET_SSE4_1) + if (TARGET_SSE_MOVCC_USE_BLENDV && TARGET_SSE4_1) gen = gen_sse4_1_blendvps; break; case E_V2DFmode: - if (TARGET_SSE4_1) + if (TARGET_SSE_MOVCC_USE_BLENDV && TARGET_SSE4_1) gen = gen_sse4_1_blendvpd; break; case E_SFmode: - if (TARGET_SSE4_1) + if (TARGET_SSE_MOVCC_USE_BLENDV && TARGET_SSE4_1) gen = gen_sse4_1_blendvss; break; case E_DFmode: - if (TARGET_SSE4_1) + if (TARGET_SSE_MOVCC_USE_BLENDV && TARGET_SSE4_1) gen = gen_sse4_1_blendvsd; break; case E_V8QImode: @@ -4278,7 +4278,7 @@ ix86_expand_sse_movcc (rtx dest, rtx cmp, rtx op_true, rtx op_false) case E_V4HFmode: case E_V4BFmode: case E_V2SImode: - if (TARGET_SSE4_1) + if (TARGET_SSE_MOVCC_USE_BLENDV && TARGET_SSE4_1) { gen = gen_mmx_pblendvb_v8qi; blend_mode = V8QImode; @@ -4288,14 +4288,14 @@ ix86_expand_sse_movcc (rtx dest, rtx cmp, rtx op_true, rtx op_false) case E_V2HImode: case E_V2HFmode: case E_V2BFmode: - if (TARGET_SSE4_1) + if (TARGET_SSE_MOVCC_USE_BLENDV && TARGET_SSE4_1) { gen = gen_mmx_pblendvb_v4qi; blend_mode = V4QImode; } break; case E_V2QImode: - if (TARGET_SSE4_1) + if (TARGET_SSE_MOVCC_USE_BLENDV && TARGET_SSE4_1) gen = gen_mmx_pblendvb_v2qi; break; case E_V16QImode: @@ -4305,18 +4305,18 @@ ix86_expand_sse_movcc (rtx dest, rtx cmp, rtx op_true, rtx op_false) case E_V4SImode: case E_V2DImode: case E_V1TImode: - if (TARGET_SSE4_1) + if (TARGET_SSE_MOVCC_USE_BLENDV && TARGET_SSE4_1) { gen = gen_sse4_1_pblendvb; blend_mode = V16QImode; } break; case E_V8SFmode: - if (TARGET_AVX) + if (TARGET_AVX && TARGET_SSE_MOVCC_USE_BLENDV) gen = gen_avx_blendvps256; break; case E_V4DFmode: - if (TARGET_AVX) + if (TARGET_AVX && TARGET_SSE_MOVCC_USE_BLENDV) gen = gen_avx_blendvpd256; break; case E_V32QImode: @@ -4325,7 +4325,7 @@ ix86_expand_sse_movcc (rtx dest, rtx cmp, rtx op_true, rtx op_false) case E_V16BFmode: case E_V8SImode: case E_V4DImode: - if (TARGET_AVX2) + if (TARGET_AVX2 && TARGET_SSE_MOVCC_USE_BLENDV) { gen = gen_avx2_pblendvb; blend_mode = V32QImode; diff --git a/gcc/config/i386/i386.h b/gcc/config/i386/i386.h index c1ec92ffb15..f01f31d208a 100644 --- a/gcc/config/i386/i386.h +++ b/gcc/config/i386/i386.h @@ -462,6 +462,8 @@ extern unsigned char ix86_tune_features[X86_TUNE_LAST]; ix86_tune_features[X86_TUNE_DEST_FALSE_DEP_FOR_GLC] #define TARGET_SLOW_STC ix86_tune_features[X86_TUNE_SLOW_STC] #define TARGET_USE_RCR ix86_tune_features[X86_TUNE_USE_RCR] +#define TARGET_SSE_MOVCC_USE_BLENDV \ + ix86_tune_features[X86_TUNE_SSE_MOVCC_USE_BLENDV] /* Feature tests against the various architecture variations. */ enum ix86_arch_indices { diff --git a/gcc/config/i386/x86-tune.def b/gcc/config/i386/x86-tune.def index 3d123da95f0..b815b6dc255 100644 --- a/gcc/config/i386/x86-tune.def +++ b/gcc/config/i386/x86-tune.def @@ -534,6 +534,14 @@ DEF_TUNE (X86_TUNE_AVOID_512FMA_CHAINS, "avoid_fma512_chains", m_ZNVER5) DEF_TUNE (X86_TUNE_V2DF_REDUCTION_PREFER_HADDPD, "v2df_reduction_prefer_haddpd", m_NONE) +/* X86_TUNE_SSE_MOVCC_USE_BLENDV: Prefer blendv instructions to + 3-instruction sequence (op1 & mask) | (op2 & ~mask) + for vector condition move. + For Crestmont, 4-operand vex blendv instructions come from MSROM + which is slow. */ +DEF_TUNE (X86_TUNE_SSE_MOVCC_USE_BLENDV, + "sse_movcc_use_blendv", ~m_CORE_ATOM) + /*****************************************************************************/ /* AVX instruction selection tuning (some of SSE flags affects AVX, too) */ /*****************************************************************************/ diff --git a/gcc/testsuite/gcc.target/i386/sse_movcc_use_blendv.c b/gcc/testsuite/gcc.target/i386/sse_movcc_use_blendv.c new file mode 100644 index 00000000000..ac9f1524949 --- /dev/null +++ b/gcc/testsuite/gcc.target/i386/sse_movcc_use_blendv.c @@ -0,0 +1,12 @@ +/* { dg-do compile } */ +/* { dg-options "-march=sierraforest -O2" } */ +/* { dg-final { scan-assembler-not {(?n)vp?blendv(b|ps|pd)} } } */ + +void +foo (int* a, int* b, int* __restrict c) +{ + for (int i = 0; i != 200; i++) + { + c[i] += a[i] > b[i] ? 1 : -1; + } +}