Text this: Current and future state of evaluation of large language models for medical summarization tasks.