Text this: Decoupling foreground and background with Siamese ViT networks for weakly-supervised semantic segmentation.