Text this: Bridging local and global representations for self-supervised monocular depth estimation.