Text this: Evaluating the clinical decision-making performance of large language models in clinically oriented thoracic anatomy scenarios