Identifying Unreported Links between ClinicalTrials.gov Trial Registrations and Their Published Results

Saved in:
Bibliographic Details
Title: Identifying Unreported Links between ClinicalTrials.gov Trial Registrations and Their Published Results
Language: English
Authors: Liu, Shifeng, Bourgeois, Florence T., Dunn, Adam G. (ORCID 0000-0002-1720-8209)
Source: Research Synthesis Methods. May 2022 13(3):342-352.
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 11
Publication Date: 2022
Sponsoring Agency: National Library of Medicine (DHHS/NIH)
Contract Number: R01LM012976
Document Type: Journal Articles
Information Analyses
Descriptors: Medical Research, Web Sites, Identification, Computational Linguistics, Comparative Analysis, Publications, Research Reports, Prediction, Drug Therapy, Users (Information), Databases
DOI: 10.1002/jrsm.1545
ISSN: 1759-2879
Abstract: A substantial proportion of trial registrations are not linked to corresponding published articles, limiting analyses and new tools. Our aim was to develop a method for finding articles reporting the results of trials that are registered on ClinicalTrials.gov when they do not include metadata links. We used a set of 27,280 trial registration and article pairs to train and evaluate methods for identifying missing links in both directions--from articles to registrations and from registrations to articles. We trained a classifier with six distance metrics as feature representations to rank the correct article or registration, using recall@K to evaluate performance and compare to baseline methods. When identifying links from registrations to published articles, the classifier ranked the correct article first (recall@1) among 378,048 articles in 80.8% of evaluation cases and 34.9% in the baseline method. Recall@10 was 85.1% compared to 60.7% in the baseline. When predicting links from articles to registrations, recall@1 was 83.4% for the classifier and 39.8% in the baseline. Recall@10 was 89.5% compared to 65.8% in the baseline. The proposed method improves on our baseline document similarity method to be feasible for identifying missing links in practice. Given a ClinicalTrials.gov registration, a user checking 10 ranked articles can expect to identify the matching article in at least 85% of cases, if the trial has been published. The proposed method can be used to improve the coupling of ClinicalTrials.gov and PubMed, with applications related to automating systematic review and evidence synthesis processes.
Abstractor: As Provided
Notes: https://github.com/evidence-surveillance/unreported_link_identidication
Entry Date: 2022
Accession Number: EJ1335011
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwGO-B_IjGfOaMv22Tl5_HkqAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDBKDvilcN9sXyyBIvwIBEICBmxMCbWOzydedjH_Q6Ee-Bt8dJslLs2p1nxK_4iUeL9Yvh1C2bLrj-V0viuzwvUT6u2aPxFRy_NqGL-oIfK3vXAaGFavxBNj9ajPDim5Y1ZN-YPBzw7KlvPO4d7LKOPXuuTQvDyxENeSYIJUCkDPtuNZOlTHQkPevWLxqpe9isKQzMu9liu7KhLJ8hQjjGwFUqfmQkjI3-5XFQQGn
Text:
  Availability: 1
  Value: <anid>AN0156769197;[bdct]01may.22;2022May11.07:24;v2.2.500</anid> <title id="AN0156769197-1">Identifying unreported links between ClinicalTrials.gov trial registrations and their published results </title> <p>A substantial proportion of trial registrations are not linked to corresponding published articles, limiting analyses and new tools. Our aim was to develop a method for finding articles reporting the results of trials that are registered on ClinicalTrials.gov when they do not include metadata links. We used a set of 27,280 trial registration and article pairs to train and evaluate methods for identifying missing links in both directions—from articles to registrations and from registrations to articles. We trained a classifier with six distance metrics as feature representations to rank the correct article or registration, using recall@K to evaluate performance and compare to baseline methods. When identifying links from registrations to published articles, the classifier ranked the correct article first (recall@1) among 378,048 articles in 80.8% of evaluation cases and 34.9% in the baseline method. Recall@10 was 85.1% compared to 60.7% in the baseline. When predicting links from articles to registrations, recall@1 was 83.4% for the classifier and 39.8% in the baseline. Recall@10 was 89.5% compared to 65.8% in the baseline. The proposed method improves on our baseline document similarity method to be feasible for identifying missing links in practice. Given a ClinicalTrials.gov registration, a user checking 10 ranked articles can expect to identify the matching article in at least 85% of cases, if the trial has been published. The proposed method can be used to improve the coupling of ClinicalTrials.gov and PubMed, with applications related to automating systematic review and evidence synthesis processes.</p> <p>Keywords: clinical trials; information retrieval; trial registration</p> <hd id="AN0156769197-2">Highlights</hd> <p></p> <hd id="AN0156769197-3">What is already known?</hd> <p></p> <ulist> <item> Many trial registrations do not include metadata links to their published results, limiting their use in meta‐research and adding barriers to efficient systematic review processes.</item> <p></p> <item> Systematic reviewers often do not incorporate registry information linked to included studies, including changes in outcomes, or summary results that may differ from what is reported in articles.</item> </ulist> <hd id="AN0156769197-4">What is new?</hd> <p></p> <ulist> <item> A new approach to learning combinations of document similarity measures for ranking candidate articles or registrations substantially improves on a previous method.</item> <p></p> <item> In a manually curated set where links between registrations and articles were missing, the new approach ranked the correct article first for nearly four of every five trial registrations.</item> </ulist> <hd id="AN0156769197-5">Potential impact</hd> <p></p> <ulist> <item> The proposed method could be used to more quickly connect trial registrations and the articles reporting their results to augment systematic review data extraction.</item> <p></p> <item> The proposed method could be used to construct more complete datasets for use in meta‐research and tools for automating evidence synthesis methods.</item> </ulist> <hd id="AN0156769197-6">INTRODUCTION</hd> <p>Prospective trial registration enables tracking and monitoring of trial activity, including timely reporting of trial results and adherence to trial protocols. ClinicalTrials.gov is the largest individual trial registry and has grown substantially following mandated use by journals, funding organisations, and regulatory agencies.1–3ClinicalTrials.gov has been used as the basis for studies investigating publication and outcome reporting biases,4–7 and data from the registry have been repurposed in novel applications, such as evidence synthesis and safety assessments.8,9</p> <p>Connecting trial registrations to published articles reporting results of trials is incomplete. Up to half of all registered trials remain unpublished in the 2 years after trial completion.10 Among the set of trials that are registered on ClinicalTrials.gov and published, only around half will have a National Clinical Trial (NCT) Number included in the abstract or metadata on PubMed.11 Of the remainder, some will include the NCT Number in the full text article while others will not have any details of the trial registration associated with the article.12</p> <p>Improved linking of trial registrations and published articles would improve surveillance efforts, provide more robust data as a basis for developing new tools, and increase the efficiency of systematic reviews. In one analysis, ClinicalTrials.gov trial records were used to identify organisations that failed to report trial results,13 but in many cases the results were in fact published and had not been linked to the registration, such that automated analyses obscured the true rates of publication.11 There are new methods and tools that rely on linking registrations and articles as training data in machine learning methods,14,15 and these could be improved by more complete and robust data. The quality and efficiency of systematic reviews could be improved if it were easier for reviewers to quickly identify articles from registrations or registrations from articles, making it easier to reconcile more complete information about trial designs and results.</p> <p>A previous study used a relatively simple document similarity approach to examine methods for ranking articles relative to clinical trial registrations to identify missing links.12 The method ranked the correct article first for 40% of registrations, and manually checking the top 50 ranked articles identified 86% of the articles. This method is useful in identifying missing links, but improvement is needed to meaningfully apply the tool in evidence synthesis and other tasks.</p> <p>Our aim was to evaluate a new method for identifying missing links between ClinicalTrials.gov trial registrations and articles in PubMed.</p> <hd id="AN0156769197-7">METHODS</hd> <p>The study involved experiments in training and evaluating classifiers. Data were obtained from a large cross‐sectional set of known links between articles and registrations, and from a manually curated set of registrations without metadata links to published articles. Features used to train models were derived from a set of distance measures between the content extracted from the free text of the trial registrations and the article titles and abstracts.</p> <hd id="AN0156769197-8">Study data</hd> <p>Trial registrations were extracted from ClinicalTrials.gov via an interface built to capture links between systematic reviews and trial registrations.16 Trial registrations were included in the analysis if they were registered on or after 1 October 2007, and were labelled as completed. The date corresponds to the expanded registration requirements for studies registered with ClinicalTrials.gov.3 On PubMed, articles were included if they were labelled as a clinical trial, controlled clinical trial, or randomised controlled trial by publication type, did not have systematic review or other review type as a publication type, and were published in 2007 or later. The registrations and articles were identified and collected on 29 September 2020.</p> <p>Data extracted from ClinicalTrials.gov registrations included the brief titles, official titles, brief summaries, and descriptions from the trial registrations. For PubMed articles, the title and abstract were extracted. Where PubMed articles included NCT Numbers in their abstract or metadata, these were assumed to represent a link from a published clinical trial to a registration, and published articles were expected to have publication dates after the trial registration date. A small proportion of PubMed articles included links to more than one trial.</p> <p>Text data were tokenized using regular expression, with punctuation, stopwords, and single characters removed. The text was constructed as two sets of processed tokens including unigram tokens and bigram tokens. Each set was transformed into representations including a binary representation, a token frequency representation, and a Token Frequency Inverse Document Frequency (TF‐IDF) representation. To avoid the impact of various lengths of the text, we applied the normalised token frequency representation with the number of tokens capturing the percentage of tokens in the text. The TF‐IDF representation is given by the product of the token frequency in the text, and the inverse document frequency, which was the logarithmically scaled inverse fraction of the documents that contained the token. Inverse document frequency was calculated separately for ClinicalTrials.gov text and PubMed text. The smoothed version of inverse document frequency was used to avoid zero‐division and zero‐multiplication problems.</p> <p>The study data included 128,480 ClinicalTrials.gov documents and 378,048 PubMed documents, with 403,496 unique unigram and 16,790,509 unique bigram tokens. The most common unigram token was <emph>study</emph> (appearing in 89.82% of registrations and 71.4% of article titles and abstracts). There were 164,551 unigrams and 10,578,070 bigrams that appeared only once in registrations and once in article titles and abstracts. From this set there were 27,280 trial registrations linked via metadata to 27,280 PubMed articles by selecting the articles published closest to the registration completion data. The frequency distribution of tokens between this subset and the rest of the corpus was similar.</p> <p>To simulate a realistic scenario with unknown links, a testing dataset was manually curated with 90 registrations that had unreported links to published articles, identified from 200 registrations with trial completion dates between 1 January 2007 and 31 December 2015.12 None of the 200 registrations had reported links from articles at the time of the search. PubMed was searched manually to identify articles reporting the results of the 200 registered trials, using a search strategy common to studies examining outcome‐reporting biases.11 The search used study design information, investigator names, locations, and other identifying features to search PubMed for the matching articles. Other information used to confirm a match included the number of participants, design and length of the study, and information on dates and location of the trial. Where multiple articles were identified as reporting results, the earliest article published after trial completion was selected.</p> <hd id="AN0156769197-9">Model construction and training</hd> <p>The problem was constructed as a binary classification task, where for a given trial registration the aim was to assign a 1 to the correct matching article and 0 to all other articles. The reverse classifier was also constructed, where for a given trial article, the aim was to assign a 1 to the correct registration and 0 to all other registrations. As an extremely unbalanced problem, this introduced challenges in terms of measuring performance and in constructing useful training examples. Classifiers that return likelihoods or scores can be used to rank articles or trials, defining an order for screening to identify matched pairs.</p> <p>The features used in model training were distance measures between the different representations. These included cosine distance between token frequency and TF‐IDF representations, and Otsuka‐Ochiai similarity between binary representations. In combination with two representations (unigram and bigram), this provided six features.</p> <p>The six distance features are: (a) cosine distance between token frequency representations of unigram tokens, (b) cosine distance between token frequency representations of bigram tokens, (c) cosine distance between TF‐IDF representations of unigram tokens, (d) cosine distance between TF‐IDF representations of bigram tokens, (e) Otsuka‐Ochiai similarity between binary representation of unigram tokens (whether a unigram token appears), and (f) Otsuka‐Ochiai similarity between binary representations of bigram tokens (whether a bigram token appears).</p> <p>The random forest method produces an ensemble model containing decision trees. Advantages of using random forest models include avoiding overfitting and the interpretability of the results, indicating which of the distance measures are most useful in combination where there are relatively few features overall.</p> <p>From the set of 27,280 unique pairs of trials and articles, 1000 were selected at random as a set of positive examples for use in training. The sampling was used to reduce the computational cost of calculating distances between every pair in the two corpora. Using all possible irrelevant links would dominate the training dataset and mislead the classifier in predicting the relation between the given trial registration and article. Instead, for each positive example, five irrelevant links were uniformly randomly sampled, for a total of 6000 examples in the training dataset (Figure 1).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01may22/jrsm1545-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1545-fig-0001.jpg" title="1 Data flow through the training and evaluation process. PubMed identifiers and NCT Numbers in italics are examples of articles and trial registrations [Colour figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <hd id="AN0156769197-11">Experiments and evaluation metrics</hd> <p>To evaluate the performance of a model, we identified the link between a trial registration and article from the set of all 378,048 articles using the combination of distance measures determined by the model. We also performed the procedure in reverse, identifying the link between an article and registration from the set of all 128,480 registrations. The model produces a likelihood score, which can be used to rank all articles (or registrations) from most likely to least likely. In general, we did not perform any sophisticated optimisation of the random forest model, including using singleton models for evaluation rather than combining models and setting most hyperparameters as default instead of performing grid‐search or cross‐validation for hyperparameter selection.</p> <p>Performance of the classifiers was evaluated using two measures. Median rank is the median value of the distribution of ranks at which the correct matching candidates are found across a set of tested examples. Recall@K for trial registrations is the proportion of matching trial articles that would be found after checking the top K ranked articles per trial registration (and the top K ranked registrations per article). We reported Recall@1 (the matching candidate was top ranked) and Recall@10, as well as illustrating Recall@K in figures.</p> <p>The computation efficiency of the classifiers is also evaluated. We calculate the computation time of both feature representation construction and classifier inference per instance in either direction. Two sets of experiments were undertaken. In the first, we trained and evaluated classifiers on set of mined links, and in the second, we evaluated the trained classifiers on a manually curated dataset to determine how the classifiers performed in a more realistic scenario with missing links. Each experiment was run five times and initialized with different random seeds, and results were reported as averaged evaluation metrics with corresponding standard deviations. Note that the evaluation using the manually curated dataset was performed for a set of registrations that did not have any metadata link to any published article. To determine the performance of the method in this set, the rank can only be calculated for those where unlinked articles were found through a standard approach for identifying missing links.11</p> <p>A manual error analysis was then used to examine in detail the examples of registration and article pairs that were consistently ranked poorly in either direction. Post hoc analyses included an investigation of the importance of each of the distances as features in the model and an error analysis. To investigate the importance of the features, we used the normalised average weighting of the features across all evaluations. Feature importance weights can vary between 0 and 1, and higher values indicate which of the distance measures contributed more to the model in terms of its importance. The error analysis included an investigation of the conditions and interventions in the articles or registrations that were ranked higher than the correct match.</p> <p>All methods and experiments were constructed in Python using NumPy, scikit‐learn, and Scipy libraries. Code for the methods and experiments are available on GitHub (https://github.com/evidence-surveillance/unreported%5flink%5fidentidication). The dataset for the methods and experiments is available on Harvard Dataverse (https://doi.org/10.7910/DVN/MEROWG).</p> <hd id="AN0156769197-12">RESULTS</hd> <p></p> <hd id="AN0156769197-13">Data description</hd> <p>The dataset included 128,480 registrations from ClinicalTrials.gov and 378,048 articles from PubMed that met the inclusion criteria. Among the registrations, 27,280 (21.2%) had one or more links from articles with an NCT number in the metadata published in the same period.</p> <p>From the 200 additionally sampled registrations without linked articles, 90 were found to have unlinked articles in PubMed.12 Noting that none of these 200 had NCT Numbers included in the abstract or metadata, 30 of 90 with missing links had included an NCT number in the full text.</p> <p>There were 184,165 unique tokens in the registrations and 327, 966 unique tokens in the articles. After removing stopwords, the number of remaining tokens varied between 11 and 3125 (median 130) per registration and between 1 and 899 (median 165) per article.</p> <hd id="AN0156769197-14">Model evaluation</hd> <p>The model was evaluated on 26,820 article‐registration pairs. When ranking articles for a given registration, the baseline method (TF‐IDF with cosine distance)12 ranked 34.9% of the matching articles first (Recall@1), compared to 80.8% for the ranking model (Table 1). The median rank and interquartile range illustrate the distribution of the rank at which the correct article was found, so a median rank of 1 (IQR 1–1) indicates that in at least 75% of cases the correct article was ranked first. The results show that when identifying missing links to articles from a trial registration on ClinicalTrials.gov, the ranking model was able to directly identify the correct article (among 378,048 articles) for four in every five cases. Recall@K results show that the ranking model consistently outperformed the baseline method and performed poorly (ranking the correct articles outside the top 100) in approximately 5% of cases (Figure 2).</p> <p>1 TABLEThe performance of using the cosine similarity and ranking model in the test set of 26,280 linked pairs of registrations and articles</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Model</th><th align="left">Median Rank (IQR)</th><th align="left">Recall@1% (std)</th><th align="left">Recall@10% (std)</th></tr></thead><tbody valign="top"><tr><td>Rank registrations given an article</td></tr><tr><td>TF‐IDF cosine</td><td>3 (1–27)</td><td>39.8% (0.00)</td><td>65.8% (0.00)</td></tr><tr><td>Ranking model</td><td>1 (1–1)</td><td>83.4% (0.02)</td><td>89.5% (0.01)</td></tr><tr><td>Rank articles given a trial registration</td></tr><tr><td>TF‐IDF cosine</td><td>4 (1–57)</td><td>34.9% (0.00)</td><td>60.7% (0.00)</td></tr><tr><td>Ranking model</td><td>1 (1–1)</td><td>80.8% (0.03)</td><td>85.1% (0.02)</td></tr></tbody></table> </ephtml> </p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01may22/jrsm1545-fig-0002.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1545-fig-0002.jpg" title="2 Scatter plot of recall@K for the baseline method (orange) and ranking model (blue) when identifying published articles from registrations in an evaluation set of 26,280 pairs. Error bars (lines) and performance range (shading) are also illustrated [Colour figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <p>When ranking registrations relative to articles, the baseline method (TF‐IDF with cosine distance)12 ranked 39.8% of the matching articles first (Recall@1), compared to 83.4% for the ranking model (Table 1). This shows that when identifying missing links to ClinicalTrials.gov registrations from an article in PubMed, the ranking model was able to directly identify the correct registration (among 128,480 registrations) for more than four in every five cases. Recall@K results demonstrated that the ranking model consistently outperformed the baseline method and performed poorly (ranking the correct articles outside the top 100) in less than 5% of cases (Figure 3).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01may22/jrsm1545-fig-0003.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1545-fig-0003.jpg" title="3 Scatter plot of recall@K for the baseline method (orange) and ranking model (blue) when identifying trial registrations given a published trial article in an evaluation set of 26,280 pairs. Error bars (lines) and performance range (shading) are also illustrated [Colour figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <hd id="AN0156769197-17">Evaluation on manually curated dataset</hd> <p>The manually curated dataset included 90 linked pairs, and the model performance degraded slightly relative to the evaluation set (Table 2). For 77.3% of the tested registrations, the matching article was ranked first. After screening 10 candidate articles for each registration, 84.7% of the matching articles could be identified (Figure 4). For 76.7% of the articles, the correct registration was ranked first. After screening 10 candidate registrations for each tested article, 81.3% of the matching registrations could be identified (Figure 5).</p> <p>2 TABLEThe performance of using the cosine similarity and ranking model in the manually crafted dataset of 90 link pairs of registrations and articles</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Model</th><th align="left">Median Rank (IQR)</th><th align="left">Recall@1% (std)</th><th align="left">Recall@10% (std)</th></tr></thead><tbody valign="top"><tr><td>Rank registrations given an article</td></tr><tr><td>TF‐IDF cosine</td><td>3 (1–30)</td><td>37.8% (0.00)</td><td>63.3% (0.00)</td></tr><tr><td>Ranking model</td><td>1 (1–1)</td><td>76.7% (0.02)</td><td>81.3% (0.02)</td></tr><tr><td>Rank articles given a registration</td></tr><tr><td>TF‐IDF cosine</td><td>5 (1–96)</td><td>41.1% (0.00)</td><td>57.8% (0.00)</td></tr><tr><td>Ranking model</td><td>1 (1–1)</td><td>77.3% (0.01)</td><td>84.7% (0.01)</td></tr></tbody></table> </ephtml> </p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01may22/jrsm1545-fig-0004.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1545-fig-0004.jpg" title="4 Scatter plot of recall@K for the baseline method (orange) and ranking model (blue) selecting article candidates for a set of 90 registrations without metadata links to articles. Error bars (lines) and performance range (shading) are also illustrated [Colour figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/BDCT/01may22/jrsm1545-fig-0005.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jrsm1545-fig-0005.jpg" title="5 Scatter plot of recall@K for the baseline method (orange) and ranking model (blue) selecting registration candidates for 90 PubMed articles without metadata links to registrations. Error bars (lines) and performance range (shading) are also illustrated [Colour figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <hd id="AN0156769197-20">Post hoc analysis of the model and its errors</hd> <p>We examined the importance of each of the distance measures as features in the model (Table 3). When identifying articles from registrations, the TF‐IDF cosine distance over unigrams weight was 0.342, TF‐IDF cosine distance over bigrams weight was 0.256, and token frequency cosine distance over unigrams weight was 0.171. When identifying registrations from articles, TF‐IDF cosine distance over unigrams weight was 0.320, TF‐IDF cosine distance over bigrams weight was 0.278, and token frequency cosine distance over unigrams weight was 0.181. The Otsuka‐Ochiai weights were all less than 0.1 in both directions.</p> <p>3 TABLEThe feature importance of six distance features (sums to 1.0) in both ranking models</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Model</th><th align="left">Unigram features</th><th align="left">Bigram Features</th></tr><tr><th align="left">TF Cosine distance</th><th align="left">TFIDF Cosine distance</th><th align="left">Otsuka‐Ochiai Similarity</th><th align="left">TF Cosine distance</th><th align="left">TFIDF Cosine distance</th><th align="left">Otsuka‐Ochiai similarity</th></tr></thead><tbody valign="top"><tr><td>Rank registrations given an article</td><td /><td align="left" /><td align="left" /></tr><tr><td>Ranking model</td><td align="char" char=".">0.181</td><td>0.320</td><td>0.073</td><td align="char" char=".">0.129</td><td>0.278</td><td>0.019</td></tr><tr><td align="char" char=".">Rank articles given a trial registration</td><td align="left" /><td /><td /></tr><tr><td>Ranking model</td><td align="char" char=".">0.171</td><td>0.342</td><td>0.069</td><td align="char" char=".">0.124</td><td>0.256</td><td>0.038</td></tr></tbody></table> </ephtml> </p> <p>A more detailed error analysis revealed where and how the approach failed to identify matching articles or registrations close to the top ranked candidates. Across the 90 manually curated pairs, 68 ranked the correct article first for the registration and the correct registration first for the article. In the remaining 22 pairs, we examined the candidates that were ranked higher than the correct article or registration using the best‐performing model.</p> <p>The 124 candidate articles or registrations that were ranked higher than the correct registration or article across the 22 pairs were examined for how closely they matched by intervention and condition (Table 4). For 64 of the higher‐ranked candidates, the intervention and condition matched the corresponding registration or article. The condition was different for 5 of 124 candidate articles or registrations; the intervention was different for 22 of the 124 candidate articles or registrations; and none of the higher‐ranked candidates were mismatched in both intervention and condition.</p> <p>4 TABLECondition and intervention match among high‐ranked but incorrect candidates</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left" /><th align="left">Identical condition</th><th align="left">Inclusive condition</th><th align="left">Distinct condition</th><th align="left">Total</th></tr></thead><tbody valign="top"><tr><td>Identical intervention</td><td>5</td><td>5</td><td>3</td><td>13</td></tr><tr><td>Overlapping intervention</td><td>59</td><td>28</td><td>2</td><td>89</td></tr><tr><td>Distinct intervention</td><td>5</td><td>17</td><td>0</td><td>22</td></tr><tr><td>Total</td><td>69</td><td>50</td><td>5</td><td>124</td></tr></tbody></table> </ephtml> </p> <hd id="AN0156769197-21">Computational efficiency of the model in the evaluation process</hd> <p>The computation time is calculated on the mined data (the 26,820 article‐registration pairs), to provide a more stable evaluation of the timing (Table 5). All experiments were performed on a Macbook Pro with 2.3 GHz quad‐core inter core i7 processor and 16 GB memory. When ranking articles for a given registration, it took 7.84 s to build feature representations given a vectorized trial registration and all 378,048 candidate articles. The ranking model then took 0.12 s to calculate scores and rank all candidate articles. The process took on average 7.96 s in total.</p> <p>5 TABLEThe computation time of feature construction and model inference in seconds per instance</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Model</th><th align="left">Feature construction</th><th align="left">Model inference</th><th align="left">Total</th></tr></thead><tbody valign="top"><tr><td>Rank registrations given an article</td></tr><tr><td>ranking model</td><td>2.14</td><td>0.52</td><td>2.66</td></tr><tr><td align="char" char=".">Rank articles given a trial registration</td></tr><tr><td>ranking model</td><td>7.84</td><td>0.12</td><td>7.96</td></tr></tbody></table> </ephtml> </p> <p>When ranking registrations relative to articles, the feature representation construction process took 2.14 s and the model inference process took 0.52 s on average, for a total of 2.66 s per registration.</p> <hd id="AN0156769197-22">DISCUSSION</hd> <p>The results indicate that a simple ranking method can be used to support the identification of missing links between trial registrations and articles. The model trained on document‐level distance features substantially outperformed our previous approach, which used individual document similarity measures.12 The performance did not substantially degrade in the evaluation set, where examples were mined from known links between registrations and articles to a more realistic set of registrations that did not have metadata links to the articles reporting their results.</p> <p>Related methods include those that are used in systematic review automation tools,17 some of which are designed to increase the efficiency of screening by reducing the number needed to screen,18 or information extraction by annotating key sentences or phrases that relate to outcomes or risk of bias.19 To date, nearly all methods that have been developed to support systematic review processes have focused only on published articles as a source of knowledge. The methods proposed in this study extend these established approaches by enabling use of trial registrations. Systematic reviewers can leverage the rich data available in trial registrations to detect publication bias and selective reporting, and augment results extracted from articles with structured summary results from ClinicalTrials.gov, which often provide more comprehensive data than articles.5,7,20</p> <p>The proposed method can also be used to improve the connectivity of trial information across sources. Ongoing surveillance of trial activity and results reporting have not been feasible in the past because robust analyses require substantial manual effort.11</p> <p>If it were possible to easily update metadata links between trial registrations and the articles reporting their results, this might enable new forms of evidence surveillance at scale,21 such as monitoring of non‐publication rates. When combined with automatic extraction of outcome measures from full text articles, it might also be possible to find interventions or conditions where certain safety outcomes are measured during trials, but are frequently missing from reports, signalling potential risks to research integrity.</p> <p>To be useful, the proposed method needs to be implementable as a tool for use by systematic reviewers. A simple implementation would present a user with a ranked list of articles (given a registration) or registrations (given an article) and a stopping criterion—a recommendation about how many documents need to be checked for a match to be confident that a matching document does not exist. Future work in this area might consider how best to implement stopping criteria to optimise between human effort and tolerance for not identifying matched pairs.</p> <p>While computation time might be an important consideration, we show that the method does not need substantial resources and can be completed quickly. Another advantage of the approach is that computational costs can be centralised. Smart approaches to operationalising the methods as a tool for public use might include crowdsourcing and making systematic review data open source. For example, one of the time‐consuming components of the process is vectorizing the 128,480 registrations and 378,048 articles, but these only need to be done once. The distance calculations do not need to be calculated by each user of the tool and can also be updated centrally and made available in a publicly accessible database of trial information.16,22 The tool could be used to make ranked lists available for every trial with an article or a registration and the results of manual checks of the ranked lists could be shared publicly.</p> <p>This study had several limitations. We made choices about how to represent the text of the registrations and articles, as well as parameters in training; other options could have produced higher performance. However, the use of multiple distance measures as features in the model demonstrated strong performance and clearly outperformed the baseline method. The data used to train and evaluate the method used examples with known metadata links and because of mandates and resources available to larger organisations, trials that have metadata links may be more recent and reported in more structured formats. These biases in the data used might make it easier to identify matched pairs because of consistency between the registration and article. However, the degradation in the results in the manually curated set without metadata links was small, which suggests that the bias may not have a major impact on the performance of the method. Additional limitations included treating each part of the text in the same way, and not using investigator names, affiliations, or locations. Using these as special features could have increased the performance further. We did not evaluate the reduction in manual effort gained from using the tool in realistic scenarios or stopping criteria for deciding that there is no matching registration or article. These would be useful avenues for future research when implementing the tool for use in practice.</p> <hd id="AN0156769197-23">CONCLUSION</hd> <p>A substantial proportion of trial registrations in ClinicalTrials.gov have results reported in published articles but remain largely inaccessible due to missing links to articles. Models trained to rank articles given a registration were able to correctly identify missing links in most cases. The proposed method could be used to improve connectivity across sources of trial information. The methods could also be used in tools to improve the speed and integrity of systematic reviews or new forms of evidence surveillance that compare or synthesise data across trial registrations and published articles.</p> <hd id="AN0156769197-24">ACKNOWLEDGMENTS</hd> <p>We acknowledge Jason Dalmazzo for support with data access and management.</p> <hd id="AN0156769197-25">CONFLICTS OF INTEREST</hd> <p>The authors declare there is no potential conflicts of interest.</p> <hd id="AN0156769197-26">AUTHOR CONTRIBUTIONS</hd> <p>Shifeng Liu, Adam G. Dunn, Florence T. Bourgeois designed the study, Shifeng Liu performed the experiments, Shifeng Liu drafted the manuscript, Adam G. Dunn, Florence T. Bourgeois critically revised the manuscript.</p> <hd id="AN0156769197-27">DATA AVAILABILITY STATEMENT</hd> <p>The methods (in Python) used to access and process the bibliographic data used in the analysis and development of the methods are available on GitHub (https://github.com/evidence‐surveillance/unreported%5flink%5fidentidication). The dataset for the methods and experiments is available on Harvard Dataverse (https://doi.org/10.7910/DVN/MEROWG).</p> <ref id="AN0156769197-28"> <title> Footnotes </title> <blist> <bibl id="bib1" type="bt">1</bibl> <bibtext> Funding informationNational Library of Medicine, National Institutes of Health R01LM012976.</bibtext> </blist> </ref> <ref id="AN0156769197-29"> <title> REFERENCES </title> <blist> <bibtext> Zarin DA, Fain KM, Dobbins HD, Tse T, Williams RJ. 10‐year update on study results submitted to ClinicalTrials.gov. New Eng J Med. 2019 ; 381 (20): 1966 ‐ 1974.</bibtext> </blist> <blist> <bibl id="bib2" type="bt">2</bibl> <bibtext> Zarin DA, Tse T, Williams RJ, Rajakannan T. Update on trial registration 11 years after the ICMJE policy was established. New Eng J Med. 2017 ; 376 (4): 383 ‐ 391.</bibtext> </blist> <blist> <bibl id="bib3" type="bt">3</bibl> <bibtext> Zarin DA, Tse T, Williams RJ, Carr S. Trial reporting in ClinicalTrials.gov — the final rule. New Eng J Med. 2016 ; 375 (20): 1998 ‐ 2004.</bibtext> </blist> <blist> <bibl id="bib4" type="bt">4</bibl> <bibtext> Glasziou PP, Sanders S, Hoffmann T. Waste in covid‐19 research. BMJ. 2020 ; 369 : m1847.</bibtext> </blist> <blist> <bibl id="bib5" type="bt">5</bibl> <bibtext> Wong EKC, Lachance CC, Page MJ, et al. Selective reporting bias in randomised controlled trials from two network meta‐analyses: comparison of clinical trial registrations and their respective publications. BMJ Open. 2019 ; 9 (9): e031138.</bibtext> </blist> <blist> <bibl id="bib6" type="bt">6</bibl> <bibtext> Bourgeois FT, Murthy S, Mandl KD. Outcome reporting among drug trials registered in ClinicalTrials.gov. Ann Intern Med. 2010 ; 153 (3): 158 ‐ 166.</bibtext> </blist> <blist> <bibl id="bib7" type="bt">7</bibl> <bibtext> Hartung DM, Zarin DA, Guise JM, McDonagh M, Paynter R, Helfand M. Reporting discrepancies between the ClinicalTrials.gov results database and peer‐reviewed publications. Ann Intern Med. 2014 ; 160 (7): 477 ‐ 483.</bibtext> </blist> <blist> <bibl id="bib8" type="bt">8</bibl> <bibtext> Ma H, Weng C. Prediction of black box warning by mining patterns of convergent focus shift in clinical trial study populations using linked public data. J Biomed Inform. 2016 ; 60 : 132 ‐ 144.</bibtext> </blist> <blist> <bibl id="bib9" type="bt">9</bibl> <bibtext> Pradhan R, Hoaglin DC, Cornell M, Liu W, Wang V, Yu H. Automatic extraction of quantitative data from ClinicalTrials.gov to conduct meta‐analyses. J Clin Epidemiol. 2019 ; 105 : 92 ‐ 100.</bibtext> </blist> <blist> <bibtext> Dwan K, Gamble C, Williamson PR, Kirkham JJ. Systematic review of the empirical evidence of study publication bias and outcome reporting bias ‐ an updated review. PloS One. 2013 ; 8 (7): e66844.</bibtext> </blist> <blist> <bibtext> Bashir R, Bourgeois FT, Dunn AG. A systematic review of the processes used to link clinical trial registrations to their published results. Syst Rev. 2017 ; 6 (1): 123.</bibtext> </blist> <blist> <bibtext> Dunn AG, Coiera E, Bourgeois FT. Unreported links between trial registrations and published articles were identified using document similarity measures in a cross‐sectional analysis of ClinicalTrials.gov. J Clin Epidemiol. 2018 ; 95 : 94 ‐ 101.</bibtext> </blist> <blist> <bibtext> Powell‐Smith A, Goldacre B. The TrialsTracker: automated ongoing monitoring of failure to share clinical trial results by all major companies and research institutions [version 1; peer review: 2 approved]. F1000Research. 2016 ; 5 : 2629.</bibtext> </blist> <blist> <bibtext> Bashir R, Surian D, Dunn AG. The risk of conclusion change in systematic review updates can be estimated by learning from a database of published examples. J Clin Epidemiol. 2019 ; 110 : 42 ‐ 49.</bibtext> </blist> <blist> <bibtext> Surian D, Dunn AG, Orenstein L, Bashir R, Coiera E, Bourgeois FT. A shared latent space matrix factorisation method for recommending new trial evidence for systematic review updates. J Biomed Inform. 2018 ; 79 : 32 ‐ 40.</bibtext> </blist> <blist> <bibtext> Martin P, Surian D, Bashir R, Bourgeois FT, Dunn AG. Trial2rev: combining machine learning and crowd‐sourcing to create a shared space for updating systematic reviews. JAMIA Open. 2019 ; 2 (1): 15 ‐ 22.</bibtext> </blist> <blist> <bibtext> Tsafnat G, Glasziou P, Choong MK, Dunn A, Galgani F, Coiera E. Systematic review automation technologies. Sys Rev. 2014 ; 3 (1): 74.</bibtext> </blist> <blist> <bibtext> O'Mara‐Eves A, Thomas J, McNaught J, Miwa M, Ananiadou S. Using text mining for study identification in systematic reviews: a systematic review of current approaches. Syst Rev. 2015 ; 4 (1): 5.</bibtext> </blist> <blist> <bibtext> Marshall IJ, Kuiper J, Wallace BC. Automating risk of bias assessment for clinical trials. IEEE J Biomed Health Inform. 2015 ; 19 (4): 1406 ‐ 1412.</bibtext> </blist> <blist> <bibtext> Tang E, Ravaud P, Riveros C, Perrodeau E, Dechartres A. Comparison of serious adverse events posted at ClinicalTrials.gov and published in corresponding journal articles. BMC Med. 2015 ; 13 : 189.</bibtext> </blist> <blist> <bibtext> Dunn AG, Bourgeois FT. Is it time for computable evidence synthesis? J Am Med Inform Assoc. 2020 ; 27 (6): 972 ‐ 975.</bibtext> </blist> <blist> <bibtext> Zarin DA, Tse T. Sharing individual participant data (IPD) within the context of the trial reporting system (TRS). PLoS Med. 2016 ; 13 (1): e1001946.</bibtext> </blist> </ref> <aug> <p>By Shifeng Liu; Florence T. Bourgeois and Adam G. Dunn</p> <p>Reported by Author; Author; Author</p> </aug>
Header DbId: eric
DbLabel: ERIC
An: EJ1335011
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Identifying Unreported Links between ClinicalTrials.gov Trial Registrations and Their Published Results
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Liu%2C+Shifeng%22">Liu, Shifeng</searchLink><br /><searchLink fieldCode="AR" term="%22Bourgeois%2C+Florence+T%2E%22">Bourgeois, Florence T.</searchLink><br /><searchLink fieldCode="AR" term="%22Dunn%2C+Adam+G%2E%22">Dunn, Adam G.</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-1720-8209">0000-0002-1720-8209</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Research+Synthesis+Methods%22"><i>Research Synthesis Methods</i></searchLink>. May 2022 13(3):342-352.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 11
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2022
– Name: SourceSuprt
  Label: Sponsoring Agency
  Group: SrcSuprt
  Data: National Library of Medicine (DHHS/NIH)
– Name: NumberContract
  Label: Contract Number
  Group: NumCntrct
  Data: R01LM012976
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Information Analyses
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Medical+Research%22">Medical Research</searchLink><br /><searchLink fieldCode="DE" term="%22Web+Sites%22">Web Sites</searchLink><br /><searchLink fieldCode="DE" term="%22Identification%22">Identification</searchLink><br /><searchLink fieldCode="DE" term="%22Computational+Linguistics%22">Computational Linguistics</searchLink><br /><searchLink fieldCode="DE" term="%22Comparative+Analysis%22">Comparative Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Publications%22">Publications</searchLink><br /><searchLink fieldCode="DE" term="%22Research+Reports%22">Research Reports</searchLink><br /><searchLink fieldCode="DE" term="%22Prediction%22">Prediction</searchLink><br /><searchLink fieldCode="DE" term="%22Drug+Therapy%22">Drug Therapy</searchLink><br /><searchLink fieldCode="DE" term="%22Users+%28Information%29%22">Users (Information)</searchLink><br /><searchLink fieldCode="DE" term="%22Databases%22">Databases</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1002/jrsm.1545
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 1759-2879
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: A substantial proportion of trial registrations are not linked to corresponding published articles, limiting analyses and new tools. Our aim was to develop a method for finding articles reporting the results of trials that are registered on ClinicalTrials.gov when they do not include metadata links. We used a set of 27,280 trial registration and article pairs to train and evaluate methods for identifying missing links in both directions--from articles to registrations and from registrations to articles. We trained a classifier with six distance metrics as feature representations to rank the correct article or registration, using recall@K to evaluate performance and compare to baseline methods. When identifying links from registrations to published articles, the classifier ranked the correct article first (recall@1) among 378,048 articles in 80.8% of evaluation cases and 34.9% in the baseline method. Recall@10 was 85.1% compared to 60.7% in the baseline. When predicting links from articles to registrations, recall@1 was 83.4% for the classifier and 39.8% in the baseline. Recall@10 was 89.5% compared to 65.8% in the baseline. The proposed method improves on our baseline document similarity method to be feasible for identifying missing links in practice. Given a ClinicalTrials.gov registration, a user checking 10 ranked articles can expect to identify the matching article in at least 85% of cases, if the trial has been published. The proposed method can be used to improve the coupling of ClinicalTrials.gov and PubMed, with applications related to automating systematic review and evidence synthesis processes.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: Note
  Label: Notes
  Group: Note
  Data: https://github.com/evidence-surveillance/unreported_link_identidication
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2022
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1335011
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1335011
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1002/jrsm.1545
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 11
        StartPage: 342
    Subjects:
      – SubjectFull: Medical Research
        Type: general
      – SubjectFull: Web Sites
        Type: general
      – SubjectFull: Identification
        Type: general
      – SubjectFull: Computational Linguistics
        Type: general
      – SubjectFull: Comparative Analysis
        Type: general
      – SubjectFull: Publications
        Type: general
      – SubjectFull: Research Reports
        Type: general
      – SubjectFull: Prediction
        Type: general
      – SubjectFull: Drug Therapy
        Type: general
      – SubjectFull: Users (Information)
        Type: general
      – SubjectFull: Databases
        Type: general
    Titles:
      – TitleFull: Identifying Unreported Links between ClinicalTrials.gov Trial Registrations and Their Published Results
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Liu, Shifeng
      – PersonEntity:
          Name:
            NameFull: Bourgeois, Florence T.
      – PersonEntity:
          Name:
            NameFull: Dunn, Adam G.
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 05
              Type: published
              Y: 2022
          Identifiers:
            – Type: issn-print
              Value: 1759-2879
          Numbering:
            – Type: volume
              Value: 13
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: Research Synthesis Methods
              Type: main
ResultId 1