{"id":872,"date":"2019-06-24T16:00:45","date_gmt":"2019-06-24T16:00:45","guid":{"rendered":"http:\/\/replicationmarkets.com\/?p=872"},"modified":"2019-10-21T10:03:11","modified_gmt":"2019-10-21T14:03:11","slug":"what-is-high-powered-replication","status":"publish","type":"post","link":"https:\/\/replicationmarkets.org\/index.php\/2019\/06\/24\/what-is-high-powered-replication\/","title":{"rendered":"What is high quality replication?"},"content":{"rendered":"\t\t<div data-elementor-type=\"wp-post\" data-elementor-id=\"872\" class=\"elementor elementor-872\">\n\t\t\t\t\t\t<section class=\"elementor-section elementor-top-section elementor-element elementor-element-521c672 elementor-section-boxed elementor-section-height-default elementor-section-height-default\" data-id=\"521c672\" data-element_type=\"section\" data-e-type=\"section\">\n\t\t\t\t\t\t<div class=\"elementor-container elementor-column-gap-default\">\n\t\t\t\t\t<div class=\"elementor-column elementor-col-100 elementor-top-column elementor-element elementor-element-840f80c\" data-id=\"840f80c\" data-element_type=\"column\" data-e-type=\"column\">\n\t\t\t<div class=\"elementor-widget-wrap elementor-element-populated\">\n\t\t\t\t\t\t<div class=\"elementor-element elementor-element-c847543 elementor-drop-cap-yes elementor-drop-cap-view-default elementor-widget elementor-widget-text-editor\" data-id=\"c847543\" data-element_type=\"widget\" data-e-type=\"widget\" data-settings=\"{&quot;drop_cap&quot;:&quot;yes&quot;}\" data-widget_type=\"text-editor.default\">\n\t\t\t\t<div class=\"elementor-widget-container\">\n\t\t\t\t\t\t\t\t\t<p><span style=\"color: #0000ff;\"><em>Note: details section (&#8220;Actually&#8221;) updated 21-OCT-2019. Additions shown in blue, deletions in <del>strikethrough<\/del>.<\/em><\/span><\/p><p>\u00a0<\/p><p><em>In brief:<\/em> A good-faith, high-power attempt to reproduce a previously-observed finding.<\/p><p><em>Slightly longer<\/em>: A good-faith attempt to reproduce a previously-observed finding, with a sample large enough to find the effect if there.<\/p><p>We will presume you already know\u00a0<a href=\"https:\/\/replicationmarkets.com\/index.php\/2019\/06\/24\/what-is-replication\/\">what is a replication<\/a>. \u00a0So&#8230; what is \u201chigh quality\u201d? Ideal first, then actual.<\/p><h3>Ideally<\/h3><p>An <em>ideal<\/em> high-quality study would be definitive: if there is an effect, it will be found, and if not, it won&#8217;t. No single study can reach this ideal, but if we did a bunch of replications, all in different labs\u00a0we might come close. Multiple replications smooth out individual quirks or mistakes, and jointly they provide a large enough sample size that it has nearly 100% power to detect a real effect \u2013 even if it\u2019s notably smaller than the original paper claimed &#8211; as most are. That&#8217;s not practical for SCORE, but amazingly it <em>has<\/em> been done at smaller scale &#8212; see for example\u00a0the <a href=\"https:\/\/cos.io\/our-services\/research\/many-labs-2-project-overview\/\">Many Labs 2<\/a> study.\u00a0<\/p><p>Such a study would come close to measuring what we\u00a0<em>really<\/em> want to know: can we trust the claim? <a href=\"https:\/\/www.darpa.mil\/program\/systematizing-confidence-in-open-research-and-evidence\">SCORE&#8217;s program description<\/a> says:<\/p><blockquote><p>Confidence scores are quantitative measures that should enable a &#8230; consumer of SBS [Social and Behavioral Science] research to understand the degree to which a particular claim or result is likely to be reproducible or replicable.\u00b9<\/p><\/blockquote><p>This ideal definition abstracts away from the limits of any particular replication, which itself may get an unlucky draw. \u00a0Therefore, DARPA suggested SCORE forecasters consider this ideal scenario:<\/p>\t\t\t\t\t\t\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t\t\t\t<\/div>\n\t\t<\/div>\n\t\t\t\t\t<\/div>\n\t\t<\/section>\n\t\t\t\t<section class=\"elementor-section elementor-top-section elementor-element elementor-element-754251d elementor-section-boxed elementor-section-height-default elementor-section-height-default\" data-id=\"754251d\" data-element_type=\"section\" data-e-type=\"section\">\n\t\t\t\t\t\t<div class=\"elementor-container elementor-column-gap-default\">\n\t\t\t\t\t<div class=\"elementor-column elementor-col-100 elementor-top-column elementor-element elementor-element-e714058\" data-id=\"e714058\" data-element_type=\"column\" data-e-type=\"column\">\n\t\t\t<div class=\"elementor-widget-wrap elementor-element-populated\">\n\t\t\t\t\t\t<div class=\"elementor-element elementor-element-783201d elementor-widget elementor-widget-blockquote\" data-id=\"783201d\" data-element_type=\"widget\" data-e-type=\"widget\" data-widget_type=\"blockquote.default\">\n\t\t\t\t<div class=\"elementor-widget-container\">\n\t\t\t\t\t \t\t\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t\t\t\t<\/div>\n\t\t<\/div>\n\t\t\t\t\t<\/div>\n\t\t<\/section>\n\t\t\t\t<section class=\"elementor-section elementor-top-section elementor-element elementor-element-d5ce59b elementor-section-boxed elementor-section-height-default elementor-section-height-default\" data-id=\"d5ce59b\" data-element_type=\"section\" data-e-type=\"section\">\n\t\t\t\t\t\t<div class=\"elementor-container elementor-column-gap-default\">\n\t\t\t\t\t<div class=\"elementor-column elementor-col-100 elementor-top-column elementor-element elementor-element-31cc74d\" data-id=\"31cc74d\" data-element_type=\"column\" data-e-type=\"column\">\n\t\t\t<div class=\"elementor-widget-wrap elementor-element-populated\">\n\t\t\t\t\t\t<div class=\"elementor-element elementor-element-2a1a208 elementor-drop-cap-yes elementor-drop-cap-view-default elementor-widget elementor-widget-text-editor\" data-id=\"2a1a208\" data-element_type=\"widget\" data-e-type=\"widget\" data-settings=\"{&quot;drop_cap&quot;:&quot;yes&quot;}\" data-widget_type=\"text-editor.default\">\n\t\t\t\t<div class=\"elementor-widget-container\">\n\t\t\t\t\t\t\t\t\t<p><span style=\"font-weight: 400;\">If you&#8217;re now trying to specify sampling distributions for studies, stop. \u00a0This is the ideal. The point is to imagine sufficient effort that everyone would agree the result is a reliable Confidence Score.\u00a0<\/span>If all our forecasters kept this frame in mind, and forecast well, we would get the desired Confidence Scores. <span style=\"font-weight: 400;\">In fact, if we&#8217;ve idealized properly, it should be a very good estimate of how much they believe the claim is <em>true<\/em>.<\/span>\u00a0<\/p><p><span style=\"font-weight: 400;\">But full disclosure: our accuracy isn&#8217;t based on the ideal, because <em>ain\u2019t nobody got time to run 100 replications each on 150+ claims<\/em>. (At only $10K apiece, that would be $150M!) \u00a0<\/span><span style=\"font-weight: 400;\">It&#8217;s based on one high-quality replication. \u00a0The forecast may be similar, but it&#8217;s unlikely to be the same. \u00a0<\/span><\/p><p><span style=\"font-weight: 400;\">So what are you really forecasting? What is &#8220;high quality&#8221; really?\u00a0<\/span><\/p><h3><b>Actually<\/b><\/h3><p>SCORE&#8217;s TA1 team will run a single high-quality replication of each selected claim. \u00a0We assume good faith and competence, so \u201chigh quality\u201d amounts to \u201chigh power\u201d. The power must be high enough that a failed replication is informative about the truth of the claim. Simonsohn (<a href=\"http:\/\/datacolada.org\/wp-content\/uploads\/2015\/05\/Small-Telescopes-Published.pdf\">2015<\/a>) argues for 2.5x the original sample size \u2013 a good rule but sometimes wasteful or practically unreachable.<\/p><p><span style=\"font-weight: 400; color: #0000ff;\">TA1 will use the two-stage approach from SSRP, noting cases where it is not feasible.<\/span><\/p><blockquote><p><span style=\"font-family: 'Helvetica Neue'; font-size: 12px; color: #0000ff;\">The sampling strategy is similar to the one used in the original studies (students or other easily accessible adult subject pools) and are described in the Replication Reports for each replication posted at OSF (<span style=\"color: #000080;\"><a style=\"color: #000080;\" href=\"https:\/\/osf.io\/pfdyw\/\">https:\/\/osf.io\/pfdyw\/<\/a><\/span>). For sample sizes we used a two-stage procedure with 90% power to detect 75% of the original effect size at the 5% level (two-sided test) in Stage 1; if the effect was not significant in the original direction in Stage 1 a second data-collection was carried out with 90% power to detect 50% of the original effect size at the 5% level (two-sided test) in the pooled first and second stage data collection.<\/span><span style=\"font-weight: 400;\">\u00a0<\/span><\/p><\/blockquote><p><span style=\"color: #0000ff;\">Previous plan below, struck out.<\/span><\/p><p><del>Therefore the TA1 gets three estimates of replication sample size, and aims for the middle:<\/del><\/p><ul><li style=\"font-weight: 400;\"><del><b>Minimum power<\/b><span style=\"font-weight: 400;\"> = <strong>smallest<\/strong> of 95% power to detect (<\/span><i><span style=\"font-weight: 400;\">p<\/span><\/i><span style=\"font-weight: 400;\"> &lt; .05) the original effect size, 80% safeguard power for 80% power, and 2.5*N<\/span><\/del><\/li><li style=\"font-weight: 400;\"><del><b>Target power<\/b><span style=\"font-weight: 400;\"> = <strong>middle<\/strong> of 95% power to detect (<\/span><i><span style=\"font-weight: 400;\">p<\/span><\/i><span style=\"font-weight: 400;\"> &lt; .05) the original effect size, 80% safeguard power for 80% power, and 2.5*N<\/span><\/del><\/li><li style=\"font-weight: 400;\"><del><b>Aspirational power<\/b><span style=\"font-weight: 400;\"> = <strong>largest<\/strong> of 95% power to detect (<\/span><i><span style=\"font-weight: 400;\">p<\/span><\/i><span style=\"font-weight: 400;\"> &lt; .05) the original effect size, 80% safeguard power for 80% power, and 2.5 \u00d7\u00a0N<\/span><\/del><\/li><\/ul><p><span style=\"font-weight: 400;\"><del>All studies will be conducted to one of these high, but imperfect, standards, aiming for at least Target Power<\/del>.\u00b2\u00a0<\/span><span style=\"font-weight: 400;\">\u00a0<\/span><\/p><p>Should this affect your forecast? \u00a0How? See\u00a0<a href=\"https:\/\/replicationmarkets.com\/index.php\/2019\/06\/24\/what-to-do\/\">What do I forecast?<\/a><\/p>\t\t\t\t\t\t\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t\t\t\t<\/div>\n\t\t<\/div>\n\t\t\t\t\t<\/div>\n\t\t<\/section>\n\t\t\t\t<section class=\"elementor-section elementor-top-section elementor-element elementor-element-b26ea65 elementor-section-boxed elementor-section-height-default elementor-section-height-default\" data-id=\"b26ea65\" data-element_type=\"section\" data-e-type=\"section\">\n\t\t\t\t\t\t<div class=\"elementor-container elementor-column-gap-default\">\n\t\t\t\t\t<div class=\"elementor-column elementor-col-100 elementor-top-column elementor-element elementor-element-00a294e\" data-id=\"00a294e\" data-element_type=\"column\" data-e-type=\"column\">\n\t\t\t<div class=\"elementor-widget-wrap elementor-element-populated\">\n\t\t\t\t\t\t<div class=\"elementor-element elementor-element-929239d elementor-widget elementor-widget-text-editor\" data-id=\"929239d\" data-element_type=\"widget\" data-e-type=\"widget\" data-widget_type=\"text-editor.default\">\n\t\t\t\t<div class=\"elementor-widget-container\">\n\t\t\t\t\t\t\t\t\t<hr \/><h3>Extra: Where do the power estimates come from?<\/h3><p><span style=\"color: #0000ff;\">As noted above, TA1 is using the two-stage approach from SSRP, but the following may still be of interest.<\/span><\/p><ul><li style=\"font-weight: 400;\"><span style=\"font-weight: 400;\"><strong>High-power<\/strong>\u00a0(<a href=\"https:\/\/science.sciencemag.org\/content\/349\/6251\/aac4716\">Open Science Collaboration, 2015<\/a>) uses 95% of the original effect size. This is the most straightforward approach, but it presumes that the original effect size estimate is credible.<\/span><\/li><li style=\"font-weight: 400;\"><span style=\"font-weight: 400;\"><strong>Safeguard power<\/strong> (<a href=\"https:\/\/doi.org\/10.1177\/1745691614528519\">Perugini et al., 2014<\/a>) bases power estimates on the lower bound of the precision of the original estimate. This is a useful strategy for precisely measured original effects while still taking into account possible inflation of published effect sizes. It is a less productive strategy (i.e., extremely demanding sample size) for imprecisely estimated effects, particularly those with a lower bound of the confidence interval close to .05. \u00a0<\/span><\/li><li style=\"font-weight: 400;\"><span style=\"font-weight: 400;\"><strong>Small telescopes<\/strong> (<a href=\"http:\/\/datacolada.org\/wp-content\/uploads\/2015\/05\/Small-Telescopes-Published.pdf\">Simonsohn, 2015<\/a>) bases power estimates on the original sample size, providing a very simple rule of using 2.5\u00d7 the original sample. This complements safeguard power in that it is (relatively) practical for highly imprecise estimated effects with a lower bound of the confidence interval close to .05. \u00a0However, it is less productive &#8212; and counterproductive for resource management &#8212; for original studies that were precisely estimated and highly significant. It is also not applicable to nested sampling such as multi-level models common in some areas of social-behavioral research (e.g., education).<\/span><\/li><\/ul><hr \/><p>\u00b9 It says &#8220;DoD consumer&#8221; because DARPA is in the US <span style=\"text-decoration: underline;\">D<\/span>ept. <span style=\"text-decoration: underline;\">o<\/span>f <span style=\"text-decoration: underline;\">D<\/span>efense, but clearly applies generally.<\/p><p>\u00b2\u200b\u00a0<span style=\"font-weight: 400;\">A few exceptions to the minimum power may be approved \u201cif the replication serves the broader interests of representativeness, coverage of domains, and completion of sufficient numbers of replications.\u201d \u00a0\u00a0As the TA1 group is keenly aware of the importance of high power, we treat this as negligible.<\/span><\/p>\t\t\t\t\t\t\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t\t\t\t<\/div>\n\t\t<\/div>\n\t\t\t\t\t<\/div>\n\t\t<\/section>\n\t\t\t\t<\/div>\n\t\t","protected":false},"excerpt":{"rendered":"<p>What is high quality replication?<br \/>\nIn brief: A good-faith, high-power attempt to reproduce a previously-observed finding.<\/p>\n","protected":false},"author":4,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_post_was_ever_published":false},"categories":[14],"tags":[69,54,27,63,55],"class_list":["post-872","post","type-post","status-publish","format-standard","hentry","category-replication","tag-confidence","tag-meta","tag-replication","tag-score","tag-statistical-power"],"jetpack_featured_media_url":"","jetpack_sharing_enabled":true,"_links":{"self":[{"href":"https:\/\/replicationmarkets.org\/index.php\/wp-json\/wp\/v2\/posts\/872","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/replicationmarkets.org\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/replicationmarkets.org\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/replicationmarkets.org\/index.php\/wp-json\/wp\/v2\/users\/4"}],"replies":[{"embeddable":true,"href":"https:\/\/replicationmarkets.org\/index.php\/wp-json\/wp\/v2\/comments?post=872"}],"version-history":[{"count":23,"href":"https:\/\/replicationmarkets.org\/index.php\/wp-json\/wp\/v2\/posts\/872\/revisions"}],"predecessor-version":[{"id":2791,"href":"https:\/\/replicationmarkets.org\/index.php\/wp-json\/wp\/v2\/posts\/872\/revisions\/2791"}],"wp:attachment":[{"href":"https:\/\/replicationmarkets.org\/index.php\/wp-json\/wp\/v2\/media?parent=872"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/replicationmarkets.org\/index.php\/wp-json\/wp\/v2\/categories?post=872"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/replicationmarkets.org\/index.php\/wp-json\/wp\/v2\/tags?post=872"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}