Text this: A scalable multimodal framework for learning engagement recognition using three-dimensional convolutional neural networks and semi-automatic annotation.