Text this: Simultaneous instance pooling and bag representation selection approach for multiple-instance learning (MIL) using vision transformer.