Why does blue stay distinct while red and amber get confused?
Because blue mostly depends on the S cone, which stays intact in both protanopia and deuteranopia, while red and amber depend on the difference between the L and M cones, exactly the ones missing in those two conditions. A hue that depends on a cone that is still present tends to survive simulation better.
Do four categorical series fit in an accessible palette, or is that always too many?
Four can work, but only if each one also carries a second identification channel: a fixed legend position, a value label next to the bar, or a per-category icon. Relying on hue alone for four groups is the easiest scenario to collapse under any of the three simulations.
Do I need to simulate all three conditions, or is one enough?
It is worth simulating all three, because each affects a different cone and can make a different pair of colors collide: a palette can pass cleanly under tritanopia and still collapse under deuteranopia, so approving it based on a single simulation lets through problems the other one would have caught.
Does swapping red and green for a different color pair solve this for good?
It solves that specific pair, but not the general problem: any new palette still needs to go through the same simulation, because the collision depends on each color exact position in the spectrum, not on "red and green" as a fixed category. Testing case by case is still necessary.