Text this: Sound-to-Image Translation Through Direct Cross-Modal Connection Using a Convolutional–Attention Generative Model.