[InstCombine] Do not simplify lshr/shl arg if it is part of fshl rotate pattern. #73441

quic-eikansh · 2023-11-26T11:10:03Z

The fshl/fshr having first two arguments as same gets lowered to targets
specific rotate. But based on the uses, one of the arguments can get
simplified resulting in different arguments performing equivalent operation.

This patch prevents the simplification of the arguments of lshr/shl if they are
part of fshl pattern.

The matchFunnelShift function was doing pattern matching and creating the fshl/fshr instruction if needed. Moved the pattern matching code to function convertShlOrLShrToFShlOrFShr. It can be reused for other optimizations.

llvmbot · 2023-11-26T11:10:32Z

@llvm/pr-subscribers-llvm-transforms

Author: None (quic-eikansh)

Changes

The fshl/fshr having first two arguments as same gets lowered to targets
specific rotate. But based on the uses, one of the arguments can get
simplified resulting in different arguments performing equivalent operation.

This patch prevents the simplification of the arguments of lshr/shl if they are
part of fshl pattern.

Full diff: https://github.com/llvm/llvm-project/pull/73441.diff

3 Files Affected:

(modified) llvm/lib/Transforms/InstCombine/InstCombineAndOrXor.cpp (+23-13)
(modified) llvm/lib/Transforms/InstCombine/InstCombineInternal.h (+3)
(modified) llvm/lib/Transforms/InstCombine/InstCombineSimplifyDemanded.cpp (+26)

diff --git a/llvm/lib/Transforms/InstCombine/InstCombineAndOrXor.cpp b/llvm/lib/Transforms/InstCombine/InstCombineAndOrXor.cpp
index 02881109f17d29f..fd4b416ec87922f 100644
--- a/llvm/lib/Transforms/InstCombine/InstCombineAndOrXor.cpp
+++ b/llvm/lib/Transforms/InstCombine/InstCombineAndOrXor.cpp
@@ -2706,9 +2706,8 @@ Instruction *InstCombinerImpl::matchBSwapOrBitReverse(Instruction &I,
   return LastInst;
 }
 
-/// Match UB-safe variants of the funnel shift intrinsic.
-static Instruction *matchFunnelShift(Instruction &Or, InstCombinerImpl &IC,
-                                     const DominatorTree &DT) {
+std::optional<std::tuple<Intrinsic::ID, SmallVector<Value *, 3>>>
+InstCombinerImpl::convertShlOrLShrToFShlOrFShr(Instruction &Or) {
   // TODO: Can we reduce the code duplication between this and the related
   // rotate matching code under visitSelect and visitTrunc?
   unsigned Width = Or.getType()->getScalarSizeInBits();
@@ -2716,7 +2715,7 @@ static Instruction *matchFunnelShift(Instruction &Or, InstCombinerImpl &IC,
   Instruction *Or0, *Or1;
   if (!match(Or.getOperand(0), m_Instruction(Or0)) ||
       !match(Or.getOperand(1), m_Instruction(Or1)))
-    return nullptr;
+    return std::nullopt;
 
   bool IsFshl = true; // Sub on LSHR.
   SmallVector<Value *, 3> FShiftArgs;
@@ -2730,7 +2729,7 @@ static Instruction *matchFunnelShift(Instruction &Or, InstCombinerImpl &IC,
         !match(Or1,
                m_OneUse(m_LogicalShift(m_Value(ShVal1), m_Value(ShAmt1)))) ||
         Or0->getOpcode() == Or1->getOpcode())
-      return nullptr;
+      return std::nullopt;
 
     // Canonicalize to or(shl(ShVal0, ShAmt0), lshr(ShVal1, ShAmt1)).
     if (Or0->getOpcode() == BinaryOperator::LShr) {
@@ -2766,7 +2765,7 @@ static Instruction *matchFunnelShift(Instruction &Or, InstCombinerImpl &IC,
       // might remove it after this fold). This still doesn't guarantee that the
       // final codegen will match this original pattern.
       if (match(R, m_OneUse(m_Sub(m_SpecificInt(Width), m_Specific(L))))) {
-        KnownBits KnownL = IC.computeKnownBits(L, /*Depth*/ 0, &Or);
+        KnownBits KnownL = computeKnownBits(L, /*Depth*/ 0, &Or);
         return KnownL.getMaxValue().ult(Width) ? L : nullptr;
       }
 
@@ -2810,7 +2809,7 @@ static Instruction *matchFunnelShift(Instruction &Or, InstCombinerImpl &IC,
       IsFshl = false; // Sub on SHL.
     }
     if (!ShAmt)
-      return nullptr;
+      return std::nullopt;
 
     FShiftArgs = {ShVal0, ShVal1, ShAmt};
   } else if (isa<ZExtInst>(Or0) || isa<ZExtInst>(Or1)) {
@@ -2832,18 +2831,18 @@ static Instruction *matchFunnelShift(Instruction &Or, InstCombinerImpl &IC,
     const APInt *ZextHighShlAmt;
     if (!match(Or0,
                m_OneUse(m_Shl(m_Value(ZextHigh), m_APInt(ZextHighShlAmt)))))
-      return nullptr;
+      return std::nullopt;
 
     if (!match(Or1, m_ZExt(m_Value(Low))) ||
         !match(ZextHigh, m_ZExt(m_Value(High))))
-      return nullptr;
+      return std::nullopt;
 
     unsigned HighSize = High->getType()->getScalarSizeInBits();
     unsigned LowSize = Low->getType()->getScalarSizeInBits();
     // Make sure High does not overlap with Low and most significant bits of
     // High aren't shifted out.
     if (ZextHighShlAmt->ult(LowSize) || ZextHighShlAmt->ugt(Width - HighSize))
-      return nullptr;
+      return std::nullopt;
 
     for (User *U : ZextHigh->users()) {
       Value *X, *Y;
@@ -2874,11 +2873,22 @@ static Instruction *matchFunnelShift(Instruction &Or, InstCombinerImpl &IC,
   }
 
   if (FShiftArgs.empty())
-    return nullptr;
+    return std::nullopt;
 
   Intrinsic::ID IID = IsFshl ? Intrinsic::fshl : Intrinsic::fshr;
-  Function *F = Intrinsic::getDeclaration(Or.getModule(), IID, Or.getType());
-  return CallInst::Create(F, FShiftArgs);
+  return std::make_tuple(IID, FShiftArgs);
+}
+
+/// Match UB-safe variants of the funnel shift intrinsic.
+static Instruction *matchFunnelShift(Instruction &Or, InstCombinerImpl &IC,
+                                     const DominatorTree &DT) {
+  if (auto Opt = IC.convertShlOrLShrToFShlOrFShr(Or)) {
+    auto [IID, FShiftArgs] = *Opt;
+    Function *F = Intrinsic::getDeclaration(Or.getModule(), IID, Or.getType());
+    return CallInst::Create(F, FShiftArgs);
+  }
+
+  return nullptr;
 }
 
 /// Attempt to combine or(zext(x),shl(zext(y),bw/2) concat packing patterns.
diff --git a/llvm/lib/Transforms/InstCombine/InstCombineInternal.h b/llvm/lib/Transforms/InstCombine/InstCombineInternal.h
index 0bbb22be71569f6..303d02cc24fc9d3 100644
--- a/llvm/lib/Transforms/InstCombine/InstCombineInternal.h
+++ b/llvm/lib/Transforms/InstCombine/InstCombineInternal.h
@@ -236,6 +236,9 @@ class LLVM_LIBRARY_VISIBILITY InstCombinerImpl final
     return getLosslessTrunc(C, TruncTy, Instruction::SExt);
   }
 
+  std::optional<std::tuple<Intrinsic::ID, SmallVector<Value *, 3>>>
+  convertShlOrLShrToFShlOrFShr(Instruction &Or);
+
 private:
   bool annotateAnyAllocSite(CallBase &Call, const TargetLibraryInfo *TLI);
   bool isDesirableIntType(unsigned BitWidth) const;
diff --git a/llvm/lib/Transforms/InstCombine/InstCombineSimplifyDemanded.cpp b/llvm/lib/Transforms/InstCombine/InstCombineSimplifyDemanded.cpp
index fa076098d63cde5..518fc84a6cca013 100644
--- a/llvm/lib/Transforms/InstCombine/InstCombineSimplifyDemanded.cpp
+++ b/llvm/lib/Transforms/InstCombine/InstCombineSimplifyDemanded.cpp
@@ -610,6 +610,19 @@ Value *InstCombinerImpl::SimplifyDemandedUseBits(Value *V, APInt DemandedMask,
                                                     DemandedMask, Known))
             return R;
 
+      // Do not simplify if shl is part of fshl rotate pattern
+      if (I->hasOneUse()) {
+        auto *Inst = dyn_cast<Instruction>(I->user_back());
+        if (Inst && Inst->getOpcode() == BinaryOperator::Or) {
+          if (auto Opt = convertShlOrLShrToFShlOrFShr(*Inst)) {
+            auto [IID, FShiftArgs] = *Opt;
+            if ((IID == Intrinsic::fshl || IID == Intrinsic::fshr) &&
+                FShiftArgs[0] == FShiftArgs[1])
+              return nullptr;
+          }
+        }
+      }
+
       // TODO: If we only want bits that already match the signbit then we don't
       // need to shift.
 
@@ -670,6 +683,19 @@ Value *InstCombinerImpl::SimplifyDemandedUseBits(Value *V, APInt DemandedMask,
     if (match(I->getOperand(1), m_APInt(SA))) {
       uint64_t ShiftAmt = SA->getLimitedValue(BitWidth-1);
 
+      // Do not simplify if lshr is part of fshl rotate pattern
+      if (I->hasOneUse()) {
+        auto *Inst = dyn_cast<Instruction>(I->user_back());
+        if (Inst && Inst->getOpcode() == BinaryOperator::Or) {
+          if (auto Opt = convertShlOrLShrToFShlOrFShr(*Inst)) {
+            auto [IID, FShiftArgs] = *Opt;
+            if ((IID == Intrinsic::fshl || IID == Intrinsic::fshr) &&
+                FShiftArgs[0] == FShiftArgs[1])
+              return nullptr;
+          }
+        }
+      }
+
       // If we are just demanding the shifted sign bit and below, then this can
       // be treated as an ASHR in disguise.
       if (DemandedMask.countl_zero() >= ShiftAmt) {

quic-eikansh · 2023-11-26T11:55:37Z

This pull request is created to address the reviews in #66115

goldsteinn · 2023-11-27T16:39:34Z

I would expect to see test difference in the commit adding the new restriction.

quic-eikansh · 2023-11-28T08:09:13Z

I would expect to see test difference in the commit adding the new restriction.

I didn't get what you meant by this. Should I have code changes and test in the same commit?

goldsteinn · 2023-11-28T16:24:11Z

I would expect to see test difference in the commit adding the new restriction.

I didn't get what you meant by this. Should I have code changes and test in the same commit?

Yes, otherwise the commit with the impl will not be passing (and it makes review more difficult as its hard to see the proper side affects of a change).

quic-eikansh · 2023-12-05T14:23:37Z

I have squashed the changes and test into one commit. @goldsteinn

goldsteinn · 2023-12-06T23:48:27Z

I have squashed the changes and test into one commit. @goldsteinn

Thats not really ideal either... Then its difficult to review the test changes.
Typically the way we do it is two seperate commits i.e

commit 1: Tests (w/ checks before the implementation)
commit 2: Implementation (re update test checks)

That way in commit 2 we can see what actual changes the implementation is causing.

…te pattern. The fshl/fshr having first two arguments as same gets lowered to targets specific rotate. But based on the uses, one of the arguments can get simplified resulting in different arguments performing equivalent operation. This patch prevents the simplification of the arguments of lshr/shl if they are part of fshl pattern.

quic-eikansh · 2023-12-13T16:27:09Z

@goldsteinn I have made the changes as you suggested.

quic-eikansh · 2024-01-08T18:11:23Z

@goldsteinn Did you get a chance to look at it?

PR Link: llvm/llvm-project#73441

dtcxzyw · 2024-01-08T19:41:50Z

llvm/lib/Transforms/InstCombine/InstCombineAndOrXor.cpp

-/// Match UB-safe variants of the funnel shift intrinsic.
-static Instruction *matchFunnelShift(Instruction &Or, InstCombinerImpl &IC,
-                                     const DominatorTree &DT) {
+std::optional<std::tuple<Intrinsic::ID, SmallVector<Value *, 3>>>


use std::pair instead?

Do std::pair has advantage over std::tuple? I see std::tuple used in codebase even for 2 element. I have addressed other 2 reviews.

I agree that we should use std::pair instead of std::tuple if there are only two elements.

llvm/lib/Transforms/InstCombine/InstCombineAndOrXor.cpp

This patch emits `ROTL(Cond, BitWidth - Shift)` directly in `ReduceSwitchRange`. This should give better codegen because `SimplifyDemandedBits` will break the rotation patterns in the original form. See also #73441 and the IR diff https://github.com/dtcxzyw/llvm-opt-benchmark/pull/115/files. This patch should cover most of cases handled by #73441.

PR Link: llvm/llvm-project#73441

This patch emits `ROTL(Cond, BitWidth - Shift)` directly in `ReduceSwitchRange`. This should give better codegen because `SimplifyDemandedBits` will break the rotation patterns in the original form. See also llvm#73441 and the IR diff https://github.com/dtcxzyw/llvm-opt-benchmark/pull/115/files. This patch should cover most of cases handled by llvm#73441.

PR Link: llvm/llvm-project#73441

dtcxzyw

LGTM. Thanks!
Please wait for additional approval from other reviewers.

quic-eikansh · 2024-02-02T09:09:27Z

@goldsteinn @nikic

quic-eikansh · 2024-02-02T09:11:44Z

LGTM. Thanks! Please wait for additional approval from other reviewers.

Thanks

dtcxzyw · 2024-02-05T18:50:15Z

Ping.

efriedma-quic

Changes LGTM... but I'd like to land as a series of three commits; one for the refactor (6bbc776 and the fixup), one for the new tests (b896ff5), then one for the actual behavior change.

Please push separate PRs for the refactor and the test, then I'll merge everything for you.

quic-eikansh · 2024-02-20T06:18:05Z

Thanks @nikic @efriedma-quic

quic-eikansh added 2 commits November 25, 2023 06:14

[InstCombine] Refactoring matchFunnelShift (NFC)

6bbc776

The matchFunnelShift function was doing pattern matching and creating the fshl/fshr instruction if needed. Moved the pattern matching code to function convertShlOrLShrToFShlOrFShr. It can be reused for other optimizations.

Update formatting in InstCombineInternal.h

1f7dec3

quic-eikansh requested a review from nikic as a code owner November 26, 2023 11:10

llvmbot added the llvm:transforms label Nov 26, 2023

quic-eikansh mentioned this pull request Nov 27, 2023

[InstCombine] Refactoring matchFunnelShift (NFC) #73390

Closed

quic-eikansh force-pushed the fsh_match branch from 5244762 to 6eceff6 Compare November 29, 2023 17:41

quic-eikansh added 2 commits December 13, 2023 07:30

[InstCombine] Added tests to fsh.ll

b896ff5

quic-eikansh force-pushed the fsh_match branch from 6eceff6 to a2300b4 Compare December 13, 2023 15:53

dtcxzyw requested a review from goldsteinn January 8, 2024 18:19

dtcxzyw added a commit to dtcxzyw/llvm-opt-benchmark that referenced this pull request Jan 8, 2024

pre-commit: test PR73441

98becbe

PR Link: llvm/llvm-project#73441

dtcxzyw mentioned this pull request Jan 8, 2024

pre-commit: test PR73441 dtcxzyw/llvm-opt-benchmark#115

Closed

dtcxzyw requested changes Jan 8, 2024

View reviewed changes

dtcxzyw mentioned this pull request Jan 10, 2024

[SimplifyCFG] Emit rotl directly in ReduceSwitchRange #77603

Merged

dtcxzyw added a commit to dtcxzyw/llvm-opt-benchmark that referenced this pull request Jan 10, 2024

pre-commit: test PR73441

cce3e27

PR Link: llvm/llvm-project#73441

Removed DT from matchFunnelShift and added assert.

0bd3c3b

quic-eikansh requested a review from dtcxzyw February 1, 2024 18:02

dtcxzyw added a commit to dtcxzyw/llvm-opt-benchmark that referenced this pull request Feb 2, 2024

pre-commit: test PR73441

3bb9908

PR Link: llvm/llvm-project#73441

dtcxzyw approved these changes Feb 2, 2024

View reviewed changes

efriedma-quic approved these changes Feb 14, 2024

View reviewed changes

nikic closed this in 3363d23 Feb 16, 2024

nikic mentioned this pull request Mar 7, 2024

[InstCombine] Do not simplify lshr/shl arg if it is part of fshl rotate pattern. #66115

Closed

[InstCombine] Do not simplify lshr/shl arg if it is part of fshl rotate pattern. #73441

[InstCombine] Do not simplify lshr/shl arg if it is part of fshl rotate pattern. #73441

Uh oh!

Conversation

quic-eikansh commented Nov 26, 2023

Uh oh!

llvmbot commented Nov 26, 2023

Uh oh!

quic-eikansh commented Nov 26, 2023

Uh oh!

goldsteinn commented Nov 27, 2023

Uh oh!

quic-eikansh commented Nov 28, 2023

Uh oh!

goldsteinn commented Nov 28, 2023

Uh oh!

quic-eikansh commented Dec 5, 2023

Uh oh!

goldsteinn commented Dec 6, 2023

Uh oh!

quic-eikansh commented Dec 13, 2023

Uh oh!

quic-eikansh commented Jan 8, 2024

Uh oh!

dtcxzyw Jan 8, 2024

Choose a reason for hiding this comment

Uh oh!

quic-eikansh Feb 1, 2024

Choose a reason for hiding this comment

Uh oh!

nikic Feb 16, 2024

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Uh oh!

dtcxzyw left a comment

Choose a reason for hiding this comment

Uh oh!

quic-eikansh commented Feb 2, 2024

Uh oh!

quic-eikansh commented Feb 2, 2024

Uh oh!

dtcxzyw commented Feb 5, 2024

Uh oh!

efriedma-quic left a comment

Choose a reason for hiding this comment

Uh oh!

quic-eikansh commented Feb 20, 2024

Uh oh!

Uh oh!