<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>R on ouR data generation</title>
    <link>https://www.rdatagen.net/tags/r/</link>
    <description>Recent content in R on ouR data generation</description>
    <generator>Hugo</generator>
    <language>en</language>
    <managingEditor>keith.goldfeld@nyumc.org (Keith Goldfeld)</managingEditor>
    <webMaster>keith.goldfeld@nyumc.org (Keith Goldfeld)</webMaster>
    <lastBuildDate>Mon, 30 Mar 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://www.rdatagen.net/tags/r/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Same model, better shape: why centering improves MCMC</title>
      <link>https://www.rdatagen.net/post/2026-03-31-centering-binary-predictors-can-improve-bayesian-computation/</link>
      <pubDate>Mon, 30 Mar 2026 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2026-03-31-centering-binary-predictors-can-improve-bayesian-computation/</guid>
      <description>&lt;p&gt;The &lt;em&gt;Emergency departments leading the transformation of Alzheimer’s and dementia care&lt;/em&gt; (ED-LEAD) study, which I have written about in the &lt;a href=&#34;https://www.rdatagen.net/post/2024-02-20-ensuring-balance-with-a-cluster-randomized-factorial-design/&#34; target=&#34;_blank&#34;&gt;past&lt;/a&gt;, is approaching the end of its third year. This multifactorial design evaluates three independent, yet potentially synergistic, interventions aimed at improving care for persons living with dementia (PLWD) and their caregivers.&lt;/p&gt;&#xA;&lt;p&gt;To estimate intervention effects, we are using what I’ve &lt;a href=&#34;https://onlinelibrary.wiley.com/doi/full/10.1002/sim.70264&#34; target=&#34;_blank&#34;&gt;called&lt;/a&gt; the &lt;em&gt;HEx-factor model&lt;/em&gt;, a Bayesian hierarchical exchangeable factorial model. The original plan was to conduct all analyses using &lt;a href=&#34;https://mc-stan.org/&#34; target=&#34;_blank&#34;&gt;&lt;code&gt;Stan&lt;/code&gt;&lt;/a&gt;. However, we’ve run into a bit of a snafu. I’ve been working through the problem, and thought I’d share here.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Getting to the bottom of TMLE: targeting in action</title>
      <link>https://www.rdatagen.net/post/2026-03-18-getting-to-the-bottom-of-tmle-targeting-in-action/</link>
      <pubDate>Wed, 18 Mar 2026 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2026-03-18-getting-to-the-bottom-of-tmle-targeting-in-action/</guid>
      <description>&lt;p&gt;In the previous &lt;a href=&#34;https://www.rdatagen.net/post/2026-03-10-getting-to-the-bottom-of-tmle-2/&#34; target=&#34;_blank&#34;&gt;post&lt;/a&gt;, I worked my way through some key elements of TMLE theory as I try to understand how it all works. At its essence, TMLE is focused on getting the efficient influence function (EIF) to behave properly. When that happens, the estimator of the target parameter behaves as if it were based on a random sample from the true data-generating distribution.&lt;/p&gt;&#xA;&lt;p&gt;Estimating the outcome and treatment (or exposure) models is an important part of constructing the EIF, but they are treated as nuisance components and do not need to be perfectly specified. The targeting step can adjust for errors in these nuisance estimates, often recovering the desired empirical behavior of the EIF and improving the resulting estimate of the target parameter, even when one of the nuisance models is misspecified.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Getting to the bottom of TMLE: forcing the target to behave</title>
      <link>https://www.rdatagen.net/post/2026-03-10-getting-to-the-bottom-of-tmle-2/</link>
      <pubDate>Mon, 09 Mar 2026 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2026-03-10-getting-to-the-bottom-of-tmle-2/</guid>
      <description>&lt;p&gt;In the last couple of posts (&lt;a href=&#34;https://www.rdatagen.net/post/2026-02-05-getting-to-the-bottom-of-tmle-1/&#34; target=&#34;_blank&#34;&gt;starting here&lt;/a&gt;), I’ve tried to unpack some of the ideas that sit underneath TMLE: viewing parameters as functionals of a distribution, thinking about sampling as a perturbation, and understanding how influence functions describe the leading behavior of estimation error. In the second &lt;a href=&#34;https://www.rdatagen.net/post/2026-03-03-getting-to-the-bottom-of-tmle-simulating-the-orthogonality/&#34; target=&#34;_blank&#34;&gt;post&lt;/a&gt;, I showed through simulation how errors in nuisance estimation can interact with sampling variability, but typically have a smaller effect than the main sampling fluctuation itself. This brings us to the central idea behind TMLE.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Getting to the bottom of TMLE: the (almost) vanishing nuisance interaction</title>
      <link>https://www.rdatagen.net/post/2026-03-03-getting-to-the-bottom-of-tmle-simulating-the-orthogonality/</link>
      <pubDate>Mon, 02 Mar 2026 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2026-03-03-getting-to-the-bottom-of-tmle-simulating-the-orthogonality/</guid>
      <description>&lt;p&gt;In the &lt;a href=&#34;https://www.rdatagen.net/post/2026-02-05-getting-to-the-bottom-of-tmle-1/&#34; target=&#34;_blank&#34;&gt;previous post&lt;/a&gt;, I argued that understanding TMLE starts with understanding how estimation error behaves. In particular, we saw that influence functions allow us to separate sampling variability from nuisance estimation error. But something subtle happens when nuisance models are estimated rather than known. The interaction term that captures their effect on the target parameter appears to shrink as the sample size grows, sometimes quite a bit. In this post, I explore that behavior through simulation. We’ll see that the nuisance interaction does shrink (though perhaps not fast enough to ignore).&lt;/p&gt;</description>
    </item>
    <item>
      <title>Getting to the bottom of TMLE: influence functions and perturbations</title>
      <link>https://www.rdatagen.net/post/2026-02-05-getting-to-the-bottom-of-tmle-1/</link>
      <pubDate>Thu, 05 Feb 2026 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2026-02-05-getting-to-the-bottom-of-tmle-1/</guid>
      <description>&lt;p&gt;I first encountered TMLE—sometimes spelled out as &lt;em&gt;targeted maximum likelihood estimation&lt;/em&gt; or &lt;em&gt;targeted minimum-loss estimate&lt;/em&gt;—about twelve or so years ago when Mark var der Laan, one of the original developers who literally wrote the &lt;a href=&#34;https://www.google.com/books/edition/Targeted_Learning/RGnSX5aCAgQC?hl=en&#34; targt=&#34;_blank&#34;&gt;book&lt;/a&gt;, gave a talk at NYU. It sounded very cool and seemed quite revolutionary and important, but it was really challenging to follow all of the details. Following that talk, I tried to tackle some of the literature, but quickly found that it as a challenge to penetrate. What struck me most was not the algorithmic complexity (which it certainly had), but much of the language and terminology, and the underlying math.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A new simstudy function to make simulating replications easier</title>
      <link>https://www.rdatagen.net/post/2025-10-07-scenario-list-facilitating-simulation-replications/</link>
      <pubDate>Tue, 07 Oct 2025 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2025-10-07-scenario-list-facilitating-simulation-replications/</guid>
      <description>&lt;p&gt;Four years ago, I &lt;a href=&#34;https://www.rdatagen.net/post/2021-03-16-framework-for-power-analysis-using-simulation/&#34; target=&#34;_blank&#34;&gt;described&lt;/a&gt; a simple framework for organizing simulations to conduct power analyses or explore the operating characteristics of modeling approaches. In that framework, I introduced a small function &lt;code&gt;scenario_list&lt;/code&gt; that generated a list of scenarios forming the basis for simulations. I had always intended to incorporate that function into &lt;code&gt;simstudy&lt;/code&gt;, and now I have finally done so The new function is available as of version &lt;code&gt;0.9.0&lt;/code&gt;.&lt;/p&gt;&#xA;&lt;p&gt;This post offers a brief introduction to the function and concludes with a small simulation.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Planning for a 3-arm cluster randomized trial with a nested intervention and a time-to-event outcome</title>
      <link>https://www.rdatagen.net/post/2025-05-20-planning-for-a-three-arm-trial-with-a-nested-intervention/</link>
      <pubDate>Tue, 20 May 2025 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2025-05-20-planning-for-a-three-arm-trial-with-a-nested-intervention/</guid>
      <description>&lt;p&gt;A researcher recently approached me for advice on a cluster-randomized trial he is developing. He is interested in testing the effectiveness of two interventions and wondered whether a 2×2 factorial design might be the best approach.&lt;/p&gt;&#xA;&lt;p&gt;As we discussed the interventions (I’ll call them &lt;span class=&#34;math inline&#34;&gt;\(A\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(B\)&lt;/span&gt;), it became clear that &lt;span class=&#34;math inline&#34;&gt;\(A\)&lt;/span&gt; was the primary focus. Intervention &lt;span class=&#34;math inline&#34;&gt;\(B\)&lt;/span&gt; might enhance the effectiveness of &lt;span class=&#34;math inline&#34;&gt;\(A\)&lt;/span&gt;, but on its own, &lt;span class=&#34;math inline&#34;&gt;\(B\)&lt;/span&gt; was not expected to have much impact. (It’s also possible that &lt;span class=&#34;math inline&#34;&gt;\(A\)&lt;/span&gt; alone doesn’t work, but once &lt;span class=&#34;math inline&#34;&gt;\(B\)&lt;/span&gt; is in place, the combination may reap benefits.) Given this, it didn’t seem worthwhile to randomize clinics or providers to receive B alone. We agreed that a three-arm cluster-randomized trial—with (1) control, (2) &lt;span class=&#34;math inline&#34;&gt;\(A\)&lt;/span&gt; alone, and (3) &lt;span class=&#34;math inline&#34;&gt;\(A + B\)&lt;/span&gt;—would be a more efficient and relevant design.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Bayesian proportional hazards model for a stepped-wedge design</title>
      <link>https://www.rdatagen.net/post/2025-04-01-bayesian-proportional-hazards-model-for-a-stepped-wedge-design/</link>
      <pubDate>Tue, 01 Apr 2025 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2025-04-01-bayesian-proportional-hazards-model-for-a-stepped-wedge-design/</guid>
      <description>&lt;p&gt;We’ve finally reached the end of the road. This is the fifth and last post in a series building up to a Bayesian proportional hazards model for analyzing a stepped-wedge cluster-randomized trial. If you are just joining in, you may want to start at the &lt;a href=&#34;https://www.rdatagen.net/post/2025-02-11-estimating-a-bayesian-proportional-hazards-model/&#34; target=&#34;_blank&#34;&gt;beginning&lt;/a&gt;.&lt;/p&gt;&#xA;&lt;p&gt;The model presented here integrates non-linear time trends and cluster-specific random effects—elements we’ve previously explored in isolation. There’s nothing fundamentally new in this post; it brings everything together. Given that the groundwork has already been laid, I’ll keep the commentary brief and focus on providing the code.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A Bayesian proportional hazards model for a cluster randomized trial</title>
      <link>https://www.rdatagen.net/post/2025-03-25-a-bayesian-proportional-hazards-model-for-a-cluster-randomized-trial/</link>
      <pubDate>Tue, 25 Mar 2025 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2025-03-25-a-bayesian-proportional-hazards-model-for-a-cluster-randomized-trial/</guid>
      <description>&lt;p&gt;In recent posts, I &lt;a href=&#34;https://www.rdatagen.net/post/2025-02-11-estimating-a-bayesian-proportional-hazards-model/&#34; target=&#34;_blank&#34;&gt;introduced&lt;/a&gt; a Bayesian approach to proportional hazards modeling and then &lt;a href=&#34;https://www.rdatagen.net/post/2025-03-04-a-bayesian-proportional-hazards-model-with-splines/&#34; target=&#34;_blank&#34;&gt;extended&lt;/a&gt; it to incorporate a penalized spline. (There was also a third &lt;a href=&#34;https://www.rdatagen.net/post/2025-03-20-bayesian-survival-model-that-can-appropriately-handle-ties/&#34; target=&#34;_blank&#34;&gt;post&lt;/a&gt; on handling ties when multiple individuals share the same event time.) This post describes another extension: a random effect to account for clustering in a cluster randomized trial. With this in place, I’ll be ready to tackle the final step—building a model for analyzing a stepped-wedge cluster-randomized trial that incorporates both splines and site-specific random effects.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Accounting for ties in a Bayesian proportional hazards model</title>
      <link>https://www.rdatagen.net/post/2025-03-20-bayesian-survival-model-that-can-appropriately-handle-ties/</link>
      <pubDate>Thu, 20 Mar 2025 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2025-03-20-bayesian-survival-model-that-can-appropriately-handle-ties/</guid>
      <description>&lt;p&gt;Over my past few posts, I’ve been progressively building towards a Bayesian model for a stepped-wedge cluster randomized trial with a time-to-event outcome, where time will be modeled using a spline function. I started with a simple Cox proportional hazards model for a traditional RCT, ignoring time as a factor. In the next post, I introduced a nonlinear time effect. For the third post—one I initially thought was ready to publish—I extended the model to a cluster randomized trial without explicitly incorporating time. I was then working on the grand finale, the full model, when I ran into an issue: I couldn’t recover the effect-size parameter used to generate the data.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A Bayesian proportional hazards model with a penalized spline</title>
      <link>https://www.rdatagen.net/post/2025-03-04-a-bayesian-proportional-hazards-model-with-splines/</link>
      <pubDate>Tue, 04 Mar 2025 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2025-03-04-a-bayesian-proportional-hazards-model-with-splines/</guid>
      <description>&lt;p&gt;In my previous &lt;a href=&#34;https://www.rdatagen.net/post/2025-02-11-estimating-a-bayesian-proportional-hazards-model/&#34; target=&#34;_blank&#34;&gt;post&lt;/a&gt;, I outlined a Bayesian approach to proportional hazards modeling. This post serves as an addendum, providing code to incorporate a spline to model a time-varying hazard ratio non linearly. In a second addendum to come I will present a separate model with a site-specific random effect, essential for a cluster-randomized trial. These components lay the groundwork for analyzing a stepped-wedge cluster-randomized trial, where both splines and site-specific random effects will be integrated into a single model. I plan on describing this comprehensive model in a final post.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Estimating a Bayesian proportional hazards model</title>
      <link>https://www.rdatagen.net/post/2025-02-11-estimating-a-bayesian-proportional-hazards-model/</link>
      <pubDate>Tue, 11 Feb 2025 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2025-02-11-estimating-a-bayesian-proportional-hazards-model/</guid>
      <description>&lt;p&gt;A recent conversation with a colleague about a large &lt;a href=&#34;https://www.rdatagen.net/post/2022-12-13-modeling-the-secular-trend-in-a-stepped-wedge-design/&#34; target=&#34;_blank&#34;&gt;stepped-wedge design&lt;/a&gt; (SW-CRT) cluster randomized trial piqued my interest, because the primary outcome is time-to-event. This is not something I’ve seen before. A quick dive into the literature suggested that time-to-event outcomes are uncommon in SW-CRTs-and that the best analytic approach is not obvious. I was intrigued by how to analyze the data to estimate a hazard ratio while accounting for clustering and potential secular trends that might influence the time to the event.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Thinking about covariates in an analysis of an RCT</title>
      <link>https://www.rdatagen.net/post/2025-01-28-handling-covariates-in-an-analysis-of-an-rct/</link>
      <pubDate>Tue, 28 Jan 2025 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2025-01-28-handling-covariates-in-an-analysis-of-an-rct/</guid>
      <description>&lt;p&gt;I was recently discussing the analytic plan for a randomized controlled trial (RCT) with a clinical collaborator when she asked whether it’s appropriate to adjust for pre-specified baseline covariates. This question is so interesting because it touches on fundamental issues of inference—both causal and statistical. What is the target estimand in an RCT—that is, what effect are we actually measuring? What do we hope to learn from the specific sample recruited for the trial (i.e., how can the findings be analyzed in a way that enhances generalizability)? What underlying assumptions about replicability, resampling, and uncertainty inform the arguments for and against covariate adjustment? These are big questions, which won’t necessarily be answered here, but need to be kept in mind when thinking about the merits of covariate adjustment&lt;/p&gt;</description>
    </item>
    <item>
      <title>Can ChatGPT help construct non-trivial statistical models? An example with Bayesian &#34;random&#34; splines</title>
      <link>https://www.rdatagen.net/post/2024-10-08-can-chatgpt-help-construct-non-trivial-bayesian-models-with-cluster-specific-splines/</link>
      <pubDate>Tue, 08 Oct 2024 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2024-10-08-can-chatgpt-help-construct-non-trivial-bayesian-models-with-cluster-specific-splines/</guid>
      <description>&lt;p&gt;I’ve been curious to see how helpful ChatGPT can be for implementing relatively complicated models in &lt;code&gt;R&lt;/code&gt;. About two years ago, I &lt;a href=&#34;https://www.rdatagen.net/post/2022-11-01-modeling-secular-trend-in-crt-using-gam/&#34; target=&#34;_blank&#34;&gt;described&lt;/a&gt; a model for estimating a treatment effect in a cluster-randomized stepped wedge trial. We used a generalized additive model (GAM) with site-specific splines to account for general time trends, implemented using the &lt;code&gt;mgcv&lt;/code&gt; package. I’ve been interested in exploring a Bayesian version of this model, but hadn’t found the time to try - until I happened to pose this simple question to ChatGPT:&lt;/p&gt;</description>
    </item>
    <item>
      <title>An IV study design to estimate an effect size when randomization is not ethical</title>
      <link>https://www.rdatagen.net/post/2024-09-03-an-instrumental-variable-study-design-to-estimate-an-effect-size-when-randomization-may-not-be-ethical/</link>
      <pubDate>Tue, 03 Sep 2024 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2024-09-03-an-instrumental-variable-study-design-to-estimate-an-effect-size-when-randomization-may-not-be-ethical/</guid>
      <description>&lt;p&gt;An investigator I frequently consult with seeks to estimate the effect of a palliative care treatment protocol for patients nearing end-stage disease, compared to a more standard, though potentially overly burdensome, therapeutic approach. Ideally, we would conduct a two-arm randomized clinical trial (RCT) to create comparable groups and obtain an unbiased estimate of the intervention effect. However, in this case, it may be considered unethical to randomize patients to a non-standard protocol.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Generating binary data by specifying the relative risk, with simulations</title>
      <link>https://www.rdatagen.net/post/2024-07-02-generating-binary-data-by-specifying-relative-risk/</link>
      <pubDate>Tue, 02 Jul 2024 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2024-07-02-generating-binary-data-by-specifying-relative-risk/</guid>
      <description>&lt;p&gt;The most traditional approach for analyzing binary outcome data is logistic regression, where the estimated parameters are interpreted as log odds ratios or, if exponentiated, as odds ratios (ORs). No one other than statisticians (and maybe not even statisticians) finds the odds ratio to be a very intuitive statistic, and many feel that a risk difference or risk ratio/relative risks (RRs) are much more interpretable. Indeed, there seems to be a strong belief that readers will, more often than not, interpret odds ratios as risk ratios. This turns out to be reasonable when an event is rare. However, when the event is more prevalent, the odds ratio will diverge from the risk ratio. (Here is a &lt;a href=&#34;https://doi.org/10.1093/aje/kwh090&#34; target=&#34;_blank&#34;&gt;paper&lt;/a&gt; that discusses some of these issues in greater depth, in case you came here looking for more.)&lt;/p&gt;</description>
    </item>
    <item>
      <title>simstudy: another way to generate data from a non-standard density</title>
      <link>https://www.rdatagen.net/post/2024-06-04-simstudy-another-way-to-generate-data-from-a-non-standard-density/</link>
      <pubDate>Tue, 04 Jun 2024 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2024-06-04-simstudy-another-way-to-generate-data-from-a-non-standard-density/</guid>
      <description>&lt;p&gt;One of my goals for the &lt;code&gt;simstudy&lt;/code&gt; package is to make it as easy as possible to generate data from a wide range of data distributions. The recent &lt;a href=&#34;https://www.rdatagen.net/post/2024-05-21-simstudy-customized-distributions/&#34; target=&#34;_blank&#34;&gt;update&lt;/a&gt; created the possibility of generating data from a customized distribution specified in a user-defined function. Last week, I added two functions, &lt;code&gt;genDataDist&lt;/code&gt; and &lt;code&gt;addDataDist&lt;/code&gt;, that allow data generation from an empirical distribution defined by a vector of integers. (See &lt;a href=&#34;https://kgoldfeld.github.io/simstudy/dev/index.html&#34; target=&#34;_blank&#34;&gt;here&lt;/a&gt; for how to download latest development version.) This post provides a simple illustration of the new functionality.&lt;/p&gt;</description>
    </item>
    <item>
      <title>simstudy 0.8.0: customized distributions</title>
      <link>https://www.rdatagen.net/post/2024-05-21-simstudy-customized-distributions/</link>
      <pubDate>Tue, 21 May 2024 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2024-05-21-simstudy-customized-distributions/</guid>
      <description>&lt;p&gt;Over the past few years, a number of folks have asked if &lt;code&gt;simstudy&lt;/code&gt; accommodates customized distributions. There’s been interest in truncated, zero-inflated, or even more standard distributions that haven’t been implemented in &lt;code&gt;simstudy&lt;/code&gt;. While I’ve come up with approaches for some of the specific cases, I was never able to develop a general solution that could provide broader flexibility.&lt;/p&gt;&#xA;&lt;p&gt;This shortcoming changes with the latest version of &lt;code&gt;simstudy&lt;/code&gt;, now available on &lt;a href=&#34;https://cran.r-project.org/web/packages/simstudy/index.html&#34; target=&#34;_blank&#34;&gt;CRAN&lt;/a&gt;. Custom distributions can now be specified in &lt;code&gt;defData&lt;/code&gt; and &lt;code&gt;defDataAdd&lt;/code&gt; by setting the argument &lt;em&gt;dist&lt;/em&gt; to “custom”. To introduce the new option, I am providing a couple of examples.&lt;/p&gt;</description>
    </item>
    <item>
      <title>simstudy enhancement: specifying idiosyncratic follow-up times for longitudinal data</title>
      <link>https://www.rdatagen.net/post/2024-04-16-simstudy-update-specifying-idiosyncratic-follow-up-times-for-longitudinal-data/</link>
      <pubDate>Tue, 16 Apr 2024 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2024-04-16-simstudy-update-specifying-idiosyncratic-follow-up-times-for-longitudinal-data/</guid>
      <description>&lt;p&gt;A researcher reached out to me a few weeks ago. They were trying to generate longitudinal data that included irregularly spaced follow-up periods. The default periods generated by the function &lt;code&gt;addPeriods&lt;/code&gt; in the &lt;code&gt;simstudy&lt;/code&gt; package are &lt;span class=&#34;math inline&#34;&gt;\(\{0, 1, 2, ..., n - 1\}\)&lt;/span&gt;, where there are &lt;span class=&#34;math inline&#34;&gt;\(n\)&lt;/span&gt; total periods. However, when follow-up periods required more specificity, such as &lt;span class=&#34;math inline&#34;&gt;\(\{0, 90, 180, 365\}\)&lt;/span&gt; days from baseline, users had to manually add them. Originally, I had intended to incorporate this feature into the function, but unfortunately it slipped through the cracks. Thanks to the clear motivation provided by the researcher, I’ve implemented this enhancement. Users can now replace the default vector with their desired set of follow-up periods using the new argument &lt;em&gt;periodVec&lt;/em&gt;. This addition is available in the development version of &lt;code&gt;simstudy&lt;/code&gt; on GitHub.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Perfectly balanced treatment arm distribution in a multifactorial CRT using stratified randomization</title>
      <link>https://www.rdatagen.net/post/2024-02-20-ensuring-balance-with-a-cluster-randomized-factorial-design/</link>
      <pubDate>Tue, 20 Feb 2024 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2024-02-20-ensuring-balance-with-a-cluster-randomized-factorial-design/</guid>
      <description>&lt;p&gt;Over two years ago, I wrote a series of posts (starting &lt;a href=&#34;https://www.rdatagen.net/post/2021-09-28-analyzing-a-factorial-trial-with-a-bayesian-model/&#34; target=&#34;_blank&#34;&gt;here&lt;/a&gt;) that described possible analytic approaches for a proposed cluster-randomized trial with a factorial design. That proposal was recently funded by NIA/NIH, and now the &lt;em&gt;Emergency departments leading the transformation of Alzheimer’s and dementia care&lt;/em&gt; (ED-LEAD) trial is just getting underway. Since the trial is in its early planning phase, I am starting to think about how we will do the randomization, and I’m sharing some of those thoughts (and code) here.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A three-arm trial using two-step randomization</title>
      <link>https://www.rdatagen.net/post/2023-12-19-a-three-arm-trial-using-two-step-randomization/</link>
      <pubDate>Tue, 19 Dec 2023 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2023-12-19-a-three-arm-trial-using-two-step-randomization/</guid>
      <description>&lt;link href=&#34;https://www.rdatagen.net/post/2023-12-19-a-three-arm-trial-using-two-step-randomization/index.en_files/tabwid/tabwid.css&#34; rel=&#34;stylesheet&#34; /&gt;&#xA;&lt;script src=&#34;https://www.rdatagen.net/post/2023-12-19-a-three-arm-trial-using-two-step-randomization/index.en_files/tabwid/tabwid.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;&lt;a href=&#34;https://bartshealth-nhs.libguides.com/CDS&#34; target=&#34;_blank&#34;&gt;Clinical Decision Support&lt;/a&gt; (CDS) tools are systems created to support clinical decision-making. Health care professionals using these tools can get guidance about diagnostic and treatment options when providing care to a patient. I’m currently involved with designing a trial focused on comparing a standard CDS tool with an enhanced version (CDS+). The main goal is to directly compare patient-level outcomes for those who have been exposed to the different versions of the CDS. However, we might also be interested in comparing the basic CDS with a control arm, which would suggest some type of three-arm trial.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Creating a nice looking Table 1 with standardized mean differences</title>
      <link>https://www.rdatagen.net/post/2023-09-26-nice-looking-table-1-with-standardized-mean-difference/</link>
      <pubDate>Tue, 26 Sep 2023 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2023-09-26-nice-looking-table-1-with-standardized-mean-difference/</guid>
      <description>&lt;link href=&#34;https://www.rdatagen.net/post/2023-09-26-nice-looking-table-1-with-standardized-mean-difference/index.en_files/table1/table1_defaults.css&#34; rel=&#34;stylesheet&#34; /&gt;&#xA;&lt;link href=&#34;https://www.rdatagen.net/post/2023-09-26-nice-looking-table-1-with-standardized-mean-difference/index.en_files/tabwid/tabwid.css&#34; rel=&#34;stylesheet&#34; /&gt;&#xA;&lt;script src=&#34;https://www.rdatagen.net/post/2023-09-26-nice-looking-table-1-with-standardized-mean-difference/index.en_files/tabwid/tabwid.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;I’m in the middle of a perfect storm, winding down three randomized clinical trials (RCTs), with patient recruitment long finished and data collection all wrapped up. This means &lt;em&gt;a lot&lt;/em&gt; of data analysis, presentation prep, and paper writing (and not so much blogging). One common (and not so glamorous) thread cutting across all of these RCTs is the need to generate a &lt;strong&gt;Table 1&lt;/strong&gt;, the comparison of baseline characteristics that convinces readers that randomization worked its magic (i.e., that study groups are indeed “comparable”). My primary goal here is to provide some &lt;code&gt;R&lt;/code&gt; code to automate the generation of this table, but not before highlighting some issues related to checking for balance and pointing you to a couple of really interesting papers.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Finding logistic models to generate data with desired risk ratio, risk difference and AUC profiles</title>
      <link>https://www.rdatagen.net/post/2023-06-20-finding-coefficients-for-logistic-models-that-generate-data-with-desired-characteristics/</link>
      <pubDate>Tue, 20 Jun 2023 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2023-06-20-finding-coefficients-for-logistic-models-that-generate-data-with-desired-characteristics/</guid>
      <description>&lt;p&gt;About two years ago, someone inquired whether &lt;code&gt;simstudy&lt;/code&gt; had the functionality to generate data from a logistic model with a specific AUC. It did not, but now it does, thanks to a &lt;a href=&#34;https://bmcmedresmethodol.biomedcentral.com/articles/10.1186/s12874-023-01836-5&#34; target=&#34;_blank&#34;&gt;paper&lt;/a&gt; by Peter Austin that describes a nice algorithm to accomplish this. The paper actually describes a series of related algorithms for generating coefficients that target specific prevalence rates, risk ratios, and risk differences, in addition to the AUC. &lt;code&gt;simstudy&lt;/code&gt; has a new function &lt;code&gt;logisticCoefs&lt;/code&gt; that implements all of these. (The Austin paper also describes an additional algorithm focused on survival outcome data and hazard ratios, but that has not been implemented in &lt;code&gt;simstudy&lt;/code&gt;). This post describes the the new function and provides some simple examples.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A demo of power estimation by simulation for a cluster randomized trial with a time-to-event outcome</title>
      <link>https://www.rdatagen.net/post/2023-05-23-just-a-simple-demo-power-estimates-for-cluster-randomized-trial-with-time-to-event-outcome/</link>
      <pubDate>Tue, 23 May 2023 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2023-05-23-just-a-simple-demo-power-estimates-for-cluster-randomized-trial-with-time-to-event-outcome/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2023-05-23-just-a-simple-demo-power-estimates-for-cluster-randomized-trial-with-time-to-event-outcome/index.en_files/htmlwidgets/htmlwidgets.js&#34;&gt;&lt;/script&gt;&#xA;&lt;script src=&#34;https://www.rdatagen.net/post/2023-05-23-just-a-simple-demo-power-estimates-for-cluster-randomized-trial-with-time-to-event-outcome/index.en_files/plotly-binding/plotly.js&#34;&gt;&lt;/script&gt;&#xA;&lt;script src=&#34;https://www.rdatagen.net/post/2023-05-23-just-a-simple-demo-power-estimates-for-cluster-randomized-trial-with-time-to-event-outcome/index.en_files/typedarray/typedarray.min.js&#34;&gt;&lt;/script&gt;&#xA;&lt;script src=&#34;https://www.rdatagen.net/post/2023-05-23-just-a-simple-demo-power-estimates-for-cluster-randomized-trial-with-time-to-event-outcome/index.en_files/jquery/jquery.min.js&#34;&gt;&lt;/script&gt;&#xA;&lt;link href=&#34;https://www.rdatagen.net/post/2023-05-23-just-a-simple-demo-power-estimates-for-cluster-randomized-trial-with-time-to-event-outcome/index.en_files/crosstalk/css/crosstalk.min.css&#34; rel=&#34;stylesheet&#34; /&gt;&#xA;&lt;script src=&#34;https://www.rdatagen.net/post/2023-05-23-just-a-simple-demo-power-estimates-for-cluster-randomized-trial-with-time-to-event-outcome/index.en_files/crosstalk/js/crosstalk.min.js&#34;&gt;&lt;/script&gt;&#xA;&lt;link href=&#34;https://www.rdatagen.net/post/2023-05-23-just-a-simple-demo-power-estimates-for-cluster-randomized-trial-with-time-to-event-outcome/index.en_files/plotly-htmlwidgets-css/plotly-htmlwidgets.css&#34; rel=&#34;stylesheet&#34; /&gt;&#xA;&lt;script src=&#34;https://www.rdatagen.net/post/2023-05-23-just-a-simple-demo-power-estimates-for-cluster-randomized-trial-with-time-to-event-outcome/index.en_files/plotly-main/plotly-latest.min.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;A colleague reached out for help designing a cluster randomized trial to evaluate a clinical decision support tool for primary care physicians (PCPs), which aims to improve care for high-risk patients. The outcome will be a time-to-event measure, collected at the patient level. The unit of randomization will be the PCP, and one of the key design issues is settling on the number to randomize. Surprisingly, I’ve never been involved with a study that required a clustered survival analysis. So, this particular sample size calculation is new for me, which led to the development of simulations that I can share with you. (There are some analytic solutions to this problem, but there doesn’t seem to a consensus about the best approach to use.)&lt;/p&gt;</description>
    </item>
    <item>
      <title>Generating variable cluster sizes to assess power in cluster randomized trials</title>
      <link>https://www.rdatagen.net/post/2023-04-18-generating-variable-cluster-sizes-to-assess-power-in-cluster-randomize-trials/</link>
      <pubDate>Tue, 18 Apr 2023 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2023-04-18-generating-variable-cluster-sizes-to-assess-power-in-cluster-randomize-trials/</guid>
      <description>&lt;p&gt;In recent discussions with a number of collaborators at the &lt;a href=&#34;https://impactcollaboratory.org/&#34; target=&#34;_blank&#34;&gt;NIA IMPACT Collaboratory&lt;/a&gt; about setting the sample size for a proposed cluster randomized trial, the question of &lt;em&gt;variable cluster sizes&lt;/em&gt; has come up a number of times. Given a fixed overall sample size, it is generally better (in terms of statistical power) if the sample is equally distributed across the different clusters; highly variable cluster sizes increase the standard errors of effect size estimates and reduce the ability to determine if an intervention or treatment is effective.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Implementing a one-step GEE algorithm for very large cluster sizes in R</title>
      <link>https://www.rdatagen.net/post/2023-03-21-implementing-a-1-step-gee-with-large-cluster-sizes-in-r/</link>
      <pubDate>Tue, 21 Mar 2023 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2023-03-21-implementing-a-1-step-gee-with-large-cluster-sizes-in-r/</guid>
      <description>&lt;p&gt;Very large data sets can present estimation problems for some statistical models, particularly ones that cannot avoid matrix inversion. For example, generalized estimating equations (GEE) models that are used when individual observations are correlated within groups can have severe computation challenges when the cluster sizes get too large. GEE are often used when repeated measures for an individual are collected over time; the individual is considered the cluster in this analysis. Estimation in this case is not really an issue because the cluster sizes are typically relatively small. However, if there are groups of individuals, we also need to account for correlation. Unfortunately, if these group/cluster sizes are too large - perhaps bigger than 1000 - traditional GEE estimation techniques just may not be feasible.&lt;/p&gt;</description>
    </item>
    <item>
      <title>simstudy 0.6.0 released: more flexible correlation patterns</title>
      <link>https://www.rdatagen.net/post/2023-02-21-flexible-correlation-generation-revisiting-block-matrices-for-temporal-patterns-in-simstudy/</link>
      <pubDate>Tue, 21 Feb 2023 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2023-02-21-flexible-correlation-generation-revisiting-block-matrices-for-temporal-patterns-in-simstudy/</guid>
      <description>&lt;p&gt;The new version (0.6.0) of &lt;code&gt;simstudy&lt;/code&gt; is available for download from &lt;a href=&#34;https://cran.r-project.org/web/packages/simstudy/index.html&#34; target=&#34;_blank&#34;&gt;CRAN&lt;/a&gt;. In addition to some important bug fixes, I’ve added new functionality that should make data generation with correlated data a little more flexible. In the previous &lt;a href=&#34;https://www.rdatagen.net/post/2023-02-14-flexible-correlation-generation-an-update-to-gencorgen-in-simstudy/&#34; target=&#34;_blank&#34;&gt;post&lt;/a&gt;, I described enhancements to the function &lt;code&gt;genCorMat&lt;/code&gt;. As part of &lt;em&gt;this&lt;/em&gt; release announcement, I’m describing &lt;code&gt;blockExchangeMat&lt;/code&gt; and &lt;code&gt;blockDecayMat&lt;/code&gt;, two new functions that can be used to generate correlation matrices when there is a &lt;em&gt;temporal&lt;/em&gt; element to the data generation.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Flexible correlation generation: an update to genCorMat in simstudy</title>
      <link>https://www.rdatagen.net/post/2023-02-14-flexible-correlation-generation-an-update-to-gencorgen-in-simstudy/</link>
      <pubDate>Tue, 14 Feb 2023 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2023-02-14-flexible-correlation-generation-an-update-to-gencorgen-in-simstudy/</guid>
      <description>&lt;p&gt;I’ve been slowly working on some updates to &lt;code&gt;simstudy&lt;/code&gt;, focusing mostly on the functionality to generate correlation matrices (which can be used to simulate correlated data). Here, I’m briefly describing the function &lt;code&gt;genCorMat&lt;/code&gt;, which has been updated to facilitate the generation of correlation matrices for clusters of different sizes and with potentially different correlation coefficients.&lt;/p&gt;&#xA;&lt;p&gt;I’ll briefly describe what the existing function can currently do, and then give an idea about what the enhancements will provide.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A GAM for time trends in a stepped-wedge trial with a binary outcome</title>
      <link>https://www.rdatagen.net/post/2023-01-17-a-gam-model-for-time-trends-in-a-stepped-wedge-trial-with-a-binary-outcome/</link>
      <pubDate>Tue, 17 Jan 2023 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2023-01-17-a-gam-model-for-time-trends-in-a-stepped-wedge-trial-with-a-binary-outcome/</guid>
      <description>&lt;p&gt;In a previous &lt;a href=&#34;https://www.rdatagen.net/post/2022-12-13-modeling-the-secular-trend-in-a-stepped-wedge-design/&#34; target=&#34;_blank&#34;&gt;post&lt;/a&gt;, I described some ways one might go about analyzing data from a stepped-wedge, cluster-randomized trial using a generalized additive model (a GAM), focusing on continuous outcomes. I have spent the past few weeks developing a similar model for a binary outcome, and have started to explore model comparison and methods to evaluate goodness-of-fit. The following describes some of my thought process.&lt;/p&gt;&#xA;&lt;div id=&#34;data-generation&#34; class=&#34;section level3&#34;&gt;&#xA;&lt;h3&gt;Data generation&lt;/h3&gt;&#xA;&lt;p&gt;The data generation process I am using here follows along pretty closely with the &lt;a href=&#34;https://www.rdatagen.net/post/2022-12-13-modeling-the-secular-trend-in-a-stepped-wedge-design/&#34; target=&#34;_blank&#34;&gt;earlier post&lt;/a&gt;, except, of course, the outcome has changed from continuous to binary. In this example, I’ve increased the correlation for between-period effects because it doesn’t seem like outcomes would change substantially from period to period, particularly if the time periods themselves are relatively short. The correlation still decays over time.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Modeling the secular trend in a stepped-wedge design</title>
      <link>https://www.rdatagen.net/post/2022-12-13-modeling-the-secular-trend-in-a-stepped-wedge-design/</link>
      <pubDate>Tue, 13 Dec 2022 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2022-12-13-modeling-the-secular-trend-in-a-stepped-wedge-design/</guid>
      <description>&lt;p&gt;&lt;a href=&#34;https://www.rdatagen.net/post/2022-11-01-modeling-secular-trend-in-crt-using-gam/&#34; target=&#34;_blank&#34;&gt;Recently&lt;/a&gt; I started a discussion about modeling secular trends using flexible models in the context of cluster randomized trials. I’ve been motivated by a trial I am involved with that is using a stepped-wedge study design. The initial post focused on more standard parallel designs; here, I want to extend the discussion explicitly to the stepped-wedge design.&lt;/p&gt;&#xA;&lt;div id=&#34;the-stepped-wedge-design&#34; class=&#34;section level3&#34;&gt;&#xA;&lt;h3&gt;The stepped-wedge design&lt;/h3&gt;&#xA;&lt;p&gt;Stepped-wedge designs are a special class of cluster randomized trial where each cluster is observed in both treatment arms (as opposed to the classic parallel design where only some of the clusters receive the treatment). In what is essentially a cross-over design, each cluster transitions in a single direction from control (or pre-intervention) to intervention. I’ve written about this in a number of different contexts (for example, with respect to &lt;a href=&#34;https://www.rdatagen.net/post/alternatives-to-stepped-wedge-designs/&#34;&gt;power analysis&lt;/a&gt;, &lt;a href=&#34;https://www.rdatagen.net/post/intra-cluster-correlations-over-time/&#34;&gt;complicated ICC patterns&lt;/a&gt;, &lt;a href=&#34;https://www.rdatagen.net/post/bayes-model-to-estimate-stepped-wedge-trial-with-non-trivial-icc-structure/&#34;&gt;using Bayesian models for estimation&lt;/a&gt;, &lt;a href=&#34;https://www.rdatagen.net/post/simulating-an-open-cohort-stepped-wedge-trial/&#34;&gt;open cohorts&lt;/a&gt;, and &lt;a href=&#34;https://www.rdatagen.net/post/2021-12-07-exploring-design-effects-of-stepped-wedge-designs-with-baseline-measurements/&#34;&gt;baseline measurements to improve efficiency&lt;/a&gt;).&lt;/p&gt;</description>
    </item>
    <item>
      <title>Generating clustered data with marginal correlations</title>
      <link>https://www.rdatagen.net/post/2022-11-22-generating-cluster-data-with-marginal-correlations/</link>
      <pubDate>Tue, 22 Nov 2022 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2022-11-22-generating-cluster-data-with-marginal-correlations/</guid>
      <description>&lt;p&gt;A student is working on a project to derive an analytic solution to the problem of sample size determination in the context of cluster randomized trials and repeated individual-level measurement (something I’ve &lt;a href=&#34;https://www.rdatagen.net/post/2021-11-23-design-effects-with-baseline-measurements/&#34; target=&#34;_blank&#34;&gt;thought&lt;/a&gt; a little bit about before). Though the goal is an analytic solution, we do want confirmation with simulation. So, I was a little disheartened to discover that the routines I’d developed in &lt;code&gt;simstudy&lt;/code&gt; for this were not quite up to the task. I’ve had to quickly fix that, and the updates are available in the development version of &lt;code&gt;simstudy&lt;/code&gt;, which can be downloaded using &lt;em&gt;devtools::install_github(“kgoldfeld/simstudy”)&lt;/em&gt;. While some of the changes are under the hood, I have added a new function, &lt;code&gt;genBlockMat&lt;/code&gt;, which I’ll describe here.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Modeling the secular trend in a cluster randomized trial using very flexible models</title>
      <link>https://www.rdatagen.net/post/2022-11-01-modeling-secular-trend-in-crt-using-gam/</link>
      <pubDate>Tue, 01 Nov 2022 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2022-11-01-modeling-secular-trend-in-crt-using-gam/</guid>
      <description>&lt;p&gt;A key challenge - maybe &lt;em&gt;the&lt;/em&gt; key challenge - of a &lt;a href=&#34;https://www.rdatagen.net/post/alternatives-to-stepped-wedge-designs/&#34; target=&#34;_blank&#34;&gt;stepped wedge clinical trial design&lt;/a&gt; is the threat of confounding by time. This is a cross-over design where the unit of randomization is a group or cluster, where each cluster begins in the control state and transitions to the intervention. It is the transition point that is randomized. Since outcomes could be changing over time regardless of the intervention, it is important to model the time trends when conducting the efficacy analysis. The question is &lt;em&gt;how&lt;/em&gt; we choose to model time, and I am going to suggest that we might want to use a very flexible model, such as a cubic spline or a generalized additive model (GAM).&lt;/p&gt;</description>
    </item>
    <item>
      <title>Presenting results for multinomial logistic regression: a marginal approach using propensity scores</title>
      <link>https://www.rdatagen.net/post/2022-09-20-presenting-results-for-multinomial-logistic-regression-a-marginal-approach-using-propensity-scores/</link>
      <pubDate>Tue, 20 Sep 2022 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2022-09-20-presenting-results-for-multinomial-logistic-regression-a-marginal-approach-using-propensity-scores/</guid>
      <description>&lt;html&gt;&#xA;&lt;link rel=&#34;stylesheet&#34; href=&#34;css/style.css&#34; /&gt;&#xA;&lt;/html&gt;&#xA;&lt;p&gt;Multinomial logistic regression modeling can provide an understanding of the factors influencing an unordered, categorical outcome. For example, if we are interested in identifying individual-level characteristics associated with political parties in the United States (&lt;em&gt;Democratic&lt;/em&gt;, &lt;em&gt;Republican&lt;/em&gt;, &lt;em&gt;Libertarian&lt;/em&gt;, &lt;em&gt;Green&lt;/em&gt;), a multinomial model would be a reasonable approach to for estimating the strength of the associations. In the case of a randomized trial or epidemiological study, we might be primarily interested in the effect of a specific intervention or exposure while controlling for other covariates. Unfortunately, interpreting results from a multinomial logistic model can be a bit of a challenge, particularly when there is a large number of possible responses and covariates.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Flexible simulation in simstudy with customized distribution functions</title>
      <link>https://www.rdatagen.net/post/2022-08-30-expanding-the-possibilities-of-simulation-in-simstudy-with-customized-distribution-funcdtions/</link>
      <pubDate>Tue, 30 Aug 2022 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2022-08-30-expanding-the-possibilities-of-simulation-in-simstudy-with-customized-distribution-funcdtions/</guid>
      <description>&lt;p&gt;Really, the only problem with the &lt;code&gt;simstudy&lt;/code&gt; package (😄) is that there is a hard limit to the possible probability distributions that are available (the current count is 15 - see &lt;a href=&#34;https://kgoldfeld.github.io/simstudy/articles/simstudy.html&#34; target=&#34;_blank&#34;&gt;here&lt;/a&gt; for a complete description). However, it turns out that there is more flexibility than first meets the eye, and we can easily accommodate a limitless number as long as you are willing to provide some extra code.&lt;/p&gt;&#xA;&lt;p&gt;I am going to illustrate this with two examples, first by implementing a truncated normal distribution, and second by implementing the flexible non-linear data generating algorithm that I &lt;a href=&#34;https://www.rdatagen.net/post/2022-08-09-simulating-data-from-a-non-linear-function-by-specifying-some-points-on-the-curve/&#34; target=&#34;_blank&#34;&gt;described last time&lt;/a&gt;.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Simulating data from a non-linear function by specifying a handful of points</title>
      <link>https://www.rdatagen.net/post/2022-08-09-simulating-data-from-a-non-linear-function-by-specifying-some-points-on-the-curve/</link>
      <pubDate>Tue, 09 Aug 2022 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2022-08-09-simulating-data-from-a-non-linear-function-by-specifying-some-points-on-the-curve/</guid>
      <description>&lt;p&gt;Trying to simulate data with non-linear relationships can be frustrating, since there is not always an obvious mathematical expression that will give you the shape you are looking for. I’ve come up with a relatively simple solution for somewhat complex scenarios that only requires the specification of a few points that lie on or near the desired curve. (Clearly, if the relationships are straightforward, such as relationships that can easily be represented by quadratic or cubic polynomials, there is no need to go through all this trouble.) The translation from the set of points to the desired function and finally to the simulated data is done by leveraging generalized additive modelling (GAM) methods, and is described here.&lt;/p&gt;</description>
    </item>
    <item>
      <title>simstudy updated to version 0.5.0</title>
      <link>https://www.rdatagen.net/post/2022-07-20-simstudy-updated-to-version-0-5-0/</link>
      <pubDate>Wed, 20 Jul 2022 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2022-07-20-simstudy-updated-to-version-0-5-0/</guid>
      <description>&lt;p&gt;A new &lt;a href=&#34;https://kgoldfeld.github.io/simstudy/index.html&#34; target=&#34;_blank&#34;&gt;version&lt;/a&gt; of &lt;code&gt;simstudy&lt;/code&gt; is available on &lt;a href=&#34;https://cran.r-project.org/web/packages/simstudy/index.html&#34; target=&#34;_blank&#34;&gt;CRAN&lt;/a&gt;. There are two major enhancements and several new features. In the “major” category, I would include (1) changes to survival data generation that accommodate hazard ratios that can change over time, as well as competing risks, and (2) the addition of functions to allow users to sample from existing data sets with replacement to generate “synthetic” data will real life distribution properties. Other less monumental, but important, changes were made: updates to functions &lt;code&gt;genFormula&lt;/code&gt; and &lt;code&gt;genMarkov&lt;/code&gt;, and two added utility functions, &lt;code&gt;survGetParams&lt;/code&gt; and &lt;code&gt;survParamPlot&lt;/code&gt;. (I did describe the survival data generation functions in two recent posts, &lt;a href=&#34;https://www.rdatagen.net/post/2022-03-15-adding-competing-risks-in-survival-data-generation/&#34; target=&#34;_blank&#34;&gt;here&lt;/a&gt; and &lt;a href=&#34;https://www.rdatagen.net/post/2022-03-29-simulating-non-proportional-hazards/&#34; target=&#34;_blank&#34;&gt;here&lt;/a&gt;.)&lt;/p&gt;</description>
    </item>
    <item>
      <title>To impute or not: the case of an RCT with baseline and follow-up measurements</title>
      <link>https://www.rdatagen.net/post/2022-04-12-to-impute-or-not-the-case-of-an-rct-with-baseline-and-follow-up-measurements/</link>
      <pubDate>Tue, 12 Apr 2022 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2022-04-12-to-impute-or-not-the-case-of-an-rct-with-baseline-and-follow-up-measurements/</guid>
      <description>&lt;p&gt;Under normal conditions, conducting a randomized clinical trial is challenging. Throw in a pandemic and things like site selection, patient recruitment and patient follow-up can be particularly vexing. In any study, subjects need to be retained long enough so that outcomes can be measured; during a period when there are so many potential disruptions, this can become quite difficult. This issue of &lt;em&gt;loss to follow-up&lt;/em&gt; recently came up during a conversation among a group of researchers who were troubleshooting challenges they are all experiencing in their ongoing trials. While everyone agreed that missing outcome data is a significant issue, there was less agreement on how to handle this analytically when estimating treatment effects.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Simulating time-to-event outcomes with non-proportional hazards</title>
      <link>https://www.rdatagen.net/post/2022-03-29-simulating-non-proportional-hazards/</link>
      <pubDate>Tue, 29 Mar 2022 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2022-03-29-simulating-non-proportional-hazards/</guid>
      <description>&lt;p&gt;As I mentioned last &lt;a href=&#34;https://www.rdatagen.net/post/2022-03-15-adding-competing-risks-in-survival-data-generation/&#34;&gt;time&lt;/a&gt;, I am working on an update of &lt;code&gt;simstudy&lt;/code&gt; that will make generating survival/time-to-event data a bit more flexible. I previously presented the functionality related to &lt;a href=&#34;https://www.rdatagen.net/post/2022-03-15-adding-competing-risks-in-survival-data-generation/&#34;&gt;competing risks&lt;/a&gt;, and this time I’ll describe generating survival data that has time-dependent hazard ratios. (As I mentioned last time, if you want to try this at home, you will need the development version of &lt;code&gt;simstudy&lt;/code&gt; that you can install using &lt;strong&gt;devtools::install_github(“kgoldfeld/simstudy”)&lt;/strong&gt;.)&lt;/p&gt;</description>
    </item>
    <item>
      <title>Adding competing risks in survival data generation</title>
      <link>https://www.rdatagen.net/post/2022-03-15-adding-competing-risks-in-survival-data-generation/</link>
      <pubDate>Tue, 15 Mar 2022 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2022-03-15-adding-competing-risks-in-survival-data-generation/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2022-03-15-adding-competing-risks-in-survival-data-generation/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;I am working on an update of &lt;code&gt;simstudy&lt;/code&gt; that will make generating survival/time-to-event data a bit more flexible. There are two biggish enhancements. The first facilitates generation of competing events, and the second allows for the possibility of generating survival data that has time-dependent hazard ratios. This post focuses on the first enhancement, and a follow up will provide examples of the second. (If you want to try this at home, you will need the development version of &lt;code&gt;simstudy&lt;/code&gt;, which you can install using &lt;strong&gt;devtools::install_github(“kgoldfeld/simstudy”)&lt;/strong&gt;.)&lt;/p&gt;</description>
    </item>
    <item>
      <title>Follow-up: simstudy function for generating parameters for survival distribution</title>
      <link>https://www.rdatagen.net/post/2022-02-22-follow-up-simstudy-function-for-generating-parameters-for-survival-distribution/</link>
      <pubDate>Tue, 22 Feb 2022 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2022-02-22-follow-up-simstudy-function-for-generating-parameters-for-survival-distribution/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2022-02-22-follow-up-simstudy-function-for-generating-parameters-for-survival-distribution/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;In the &lt;a href=&#34;https://www.rdatagen.net/post/2022-02-08-simulating-survival-outcomes-setting-the-parameters-for-the-desired-distribution/&#34; target=&#34;_blank&#34;&gt;previous post&lt;/a&gt; I described how to determine the parameter values for generating a Weibull survival curve that reflects a desired distribution defined by two points along the curve. I went ahead and implemented these ideas in the development version of &lt;code&gt;simstudy 0.4.0.9000&lt;/code&gt;, expanding the idea to allow for any number of points rather than just two. This post provides a brief overview of the approach, the code, and a simple example using the parameters to generate simulated data.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Simulating survival outcomes: setting the parameters for the desired distribution</title>
      <link>https://www.rdatagen.net/post/2022-02-08-simulating-survival-outcomes-setting-the-parameters-for-the-desired-distribution/</link>
      <pubDate>Tue, 08 Feb 2022 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2022-02-08-simulating-survival-outcomes-setting-the-parameters-for-the-desired-distribution/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2022-02-08-simulating-survival-outcomes-setting-the-parameters-for-the-desired-distribution/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;The package &lt;code&gt;simstudy&lt;/code&gt; has some functions that facilitate generating survival data using an underlying Weibull distribution. Originally, I added this to the package because I thought it would be interesting to try to do, and I figured it would be useful for me someday (and hopefully some others, as well). Well, now I am working on a project that involves evaluating at least two survival-type processes that are occurring simultaneously. To get a handle on the analytic models we might use, I’ve started to try to simulate a simplified version of the data that we have.&lt;/p&gt;</description>
    </item>
    <item>
      <title>simstudy update: ordinal data generation that violates proportionality</title>
      <link>https://www.rdatagen.net/post/2022-01-25-simstudy-update-ordinal-data-without-the-proportionality-assumption/</link>
      <pubDate>Tue, 25 Jan 2022 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2022-01-25-simstudy-update-ordinal-data-without-the-proportionality-assumption/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2022-01-25-simstudy-update-ordinal-data-without-the-proportionality-assumption/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;Version 0.4.0 of &lt;code&gt;simstudy&lt;/code&gt; is now available on &lt;a href=&#34;https://cran.r-project.org/web/packages/simstudy/index.html&#34; target=&#34;_blank&#34;&gt;CRAN&lt;/a&gt; and &lt;a href=&#34;https://github.com/kgoldfeld/simstudy&#34; target=&#34;_blank&#34;&gt;GitHub&lt;/a&gt;. This update includes two enhancements (and at least one major bug fix). &lt;code&gt;genOrdCat&lt;/code&gt; now includes an argument to generate ordinal data without an assumption of cumulative proportional odds. And two new functions &lt;code&gt;defRepeat&lt;/code&gt; and &lt;code&gt;defRepeatAdd&lt;/code&gt; make it a bit easier to define multiple variables that share the same distribution assumptions.&lt;/p&gt;&#xA;&lt;div id=&#34;ordinal-data&#34; class=&#34;section level2&#34;&gt;&#xA;&lt;h2&gt;Ordinal data&lt;/h2&gt;&#xA;&lt;p&gt;In &lt;code&gt;simstudy&lt;/code&gt;, it is relatively easy to specify multinomial distributions that characterize categorical data. Order becomes relevant when the categories take on meanings related to strength of opinion or agreement (as in a Likert-type response) or frequency. A motivating example could be when a response variable takes on four possible values: (1) strongly disagree, (2) disagree, (4) agree, (5) strongly agree. There is a natural order to the response possibilities.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Including uncertainty when comparing response rates across clusters</title>
      <link>https://www.rdatagen.net/post/2022-01-18-including-uncertainty-when-comparing-response-rates-across-clusters/</link>
      <pubDate>Tue, 18 Jan 2022 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2022-01-18-including-uncertainty-when-comparing-response-rates-across-clusters/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2022-01-18-including-uncertainty-when-comparing-response-rates-across-clusters/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;Since this is a holiday weekend here in the US, I thought I would write up something relatively short and simple since I am supposed to be relaxing. A few weeks ago, someone presented me with some data that showed response rates to a survey that was conducted at about 30 different locations. The team that collected the data was interested in understanding if there were some sites that had response rates that might have been too low. To determine this, they generated a plot that looked something like this:&lt;/p&gt;</description>
    </item>
    <item>
      <title>Skeptical Bayesian priors might help minimize skepticism about subgroup analyses</title>
      <link>https://www.rdatagen.net/post/2022-01-04-reducing-the-risk-of-spurious-findings-with-bayesian-decison-rules/</link>
      <pubDate>Tue, 04 Jan 2022 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2022-01-04-reducing-the-risk-of-spurious-findings-with-bayesian-decison-rules/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2022-01-04-reducing-the-risk-of-spurious-findings-with-bayesian-decison-rules/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;Over the past couple of years, I have been working with an amazing group of investigators as part of the CONTAIN trial to study whether COVID-19 convalescent plasma (CCP) can improve the clinical status of patients hospitalized with COVID-19 and requiring noninvasive supplemental oxygen. This was a multi-site study in the US that randomized 941 patients to either CCP or a saline solution placebo. The overall &lt;a href=&#34;https://jamanetwork.com/journals/jamainternalmedicine/article-abstract/2787090&#34; target=&#34;_blank&#34;&gt;findings&lt;/a&gt; suggest that CCP did not benefit the patients who received it, but if you drill down a little deeper, the story may be more complicated than that.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Controlling Type I error in RCTs with interim looks: a Bayesian perspective</title>
      <link>https://www.rdatagen.net/post/2021-12-21-controling-type-1-error-rates-in-rcts-with-interim-looks-a-bayesian-perspective/</link>
      <pubDate>Tue, 21 Dec 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-12-21-controling-type-1-error-rates-in-rcts-with-interim-looks-a-bayesian-perspective/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-12-21-controling-type-1-error-rates-in-rcts-with-interim-looks-a-bayesian-perspective/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;Recently, a colleague submitted a paper describing the results of a Bayesian adaptive trial where the research team estimated the probability of effectiveness at various points during the trial. This trial was designed to stop as soon as the probability of effectiveness exceeded a pre-specified threshold. The journal rejected the paper on the grounds that these repeated interim looks inflated the Type I error rate, and increased the chances that any conclusions drawn from the study could have been misleading. Was this a reasonable position for the journal editors to take?&lt;/p&gt;</description>
    </item>
    <item>
      <title>Exploring design effects of stepped wedge designs with baseline measurements</title>
      <link>https://www.rdatagen.net/post/2021-12-07-exploring-design-effects-of-stepped-wedge-designs-with-baseline-measurements/</link>
      <pubDate>Tue, 07 Dec 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-12-07-exploring-design-effects-of-stepped-wedge-designs-with-baseline-measurements/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-12-07-exploring-design-effects-of-stepped-wedge-designs-with-baseline-measurements/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;In the &lt;a href=&#34;https://www.rdatagen.net/post/2021-11-23-design-effects-with-baseline-measurements/&#34; target=&#34;_blank&#34;&gt;previous post&lt;/a&gt;, I described an incipient effort that I am undertaking with two colleagues, Monica Taljaard and Fan Li, to better understand the implications for collecting baseline measurements on sample size requirements for stepped wedge cluster randomized trials. (The three of us are on the &lt;a href=&#34;https://impactcollaboratory.org/design-and-statistics-core/&#34; target=&#34;_blank&#34;&gt;Design and Statistics Core&lt;/a&gt; of the &lt;a href=&#34;https://impactcollaboratory.org/&#34; target=&#34;_blank&#34;&gt;NIA IMPACT Collaboratory&lt;/a&gt;.) In that post, I conducted a series of simulations that illustrated the design effects in parallel cluster randomized trials derived analytically in a &lt;a href=&#34;https://onlinelibrary.wiley.com/doi/full/10.1002/sim.5352&#34; target=&#34;_blank&#34;&gt;paper&lt;/a&gt; by &lt;em&gt;Teerenstra et al&lt;/em&gt;. In this post, I am extending those simulations to stepped wedge trials; the hope is that the design effects can be formally derived some point soon.&lt;/p&gt;</description>
    </item>
    <item>
      <title>The design effect of a cluster randomized trial with baseline measurements</title>
      <link>https://www.rdatagen.net/post/2021-11-23-design-effects-with-baseline-measurements/</link>
      <pubDate>Tue, 23 Nov 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-11-23-design-effects-with-baseline-measurements/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-11-23-design-effects-with-baseline-measurements/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;Is it possible to reduce the sample size requirements of a stepped wedge cluster randomized trial simply by collecting baseline information? In a trial with randomization at the individual level, it &lt;em&gt;is&lt;/em&gt; generally the case that if we are able to measure an outcome for subjects at two time periods, first at baseline and then at follow-up, we can reduce the overall sample size. But does this extend to (a) cluster randomized trials generally, and to (b) stepped wedge designs more specifically?&lt;/p&gt;</description>
    </item>
    <item>
      <title>simstudy update: adding flexibility to data generation</title>
      <link>https://www.rdatagen.net/post/2021-11-09-simstudy-0-3-0-update-summary/</link>
      <pubDate>Tue, 09 Nov 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-11-09-simstudy-0-3-0-update-summary/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-11-09-simstudy-0-3-0-update-summary/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;A new version of &lt;code&gt;simstudy&lt;/code&gt; (0.3.0) is now available on &lt;a href=&#34;https://cran.r-project.org/web/packages/simstudy/index.html&#34; target=&#34;_blank&#34;&gt;CRAN&lt;/a&gt; and on the &lt;a href=&#34;https://github.com/kgoldfeld/simstudy/releases&#34; target=&#34;_blank&#34;&gt;package website&lt;/a&gt;. Along with some less exciting bug fixes, we have added capabilities to a few existing features: double-dot variable reference, treatment assignment, and categorical data definition. These simple additions should make the data generation process a little smoother and more flexible.&lt;/p&gt;&#xA;&lt;div id=&#34;using-non-scalar-double-dot-variable-reference&#34; class=&#34;section level2&#34;&gt;&#xA;&lt;h2&gt;Using non-scalar double-dot variable reference&lt;/h2&gt;&#xA;&lt;p&gt;Double-dot notation was &lt;a href=&#34;https://www.rdatagen.net/post/simstudy-just-got-a-little-more-dynamic-version-0-2-0/&#34;&gt;introduced&lt;/a&gt; in the last version of &lt;code&gt;simstudy&lt;/code&gt; to allow data definitions to be more dynamic. Previously, the double-dot variable could only be a scalar value, and with the current version, double-dot notation is now also &lt;em&gt;array-friendly&lt;/em&gt;.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Sample size requirements for a Bayesian factorial study design</title>
      <link>https://www.rdatagen.net/post/2021-10-26-sample-size-requirements-for-a-factorial-study-design/</link>
      <pubDate>Tue, 26 Oct 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-10-26-sample-size-requirements-for-a-factorial-study-design/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-10-26-sample-size-requirements-for-a-factorial-study-design/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;How do you determine sample size when the goal of a study is not to conduct a null hypothesis test but to provide an estimate of multiple effect sizes? I needed to get a handle on this for a recent grant submission, which I’ve been writing about over the past month, &lt;a href=&#34;https://www.rdatagen.net/post/2021-09-28-analyzing-a-factorial-trial-with-a-bayesian-model/&#34; target=&#34;_blank&#34;&gt;here&lt;/a&gt; and &lt;a href=&#34;https://www.rdatagen.net/post/2021-10-12-analyzing-a-factorial-design-with-a-bayesian-shrinkage-model/&#34; target=&#34;_blank&#34;&gt;here&lt;/a&gt;. (I provide a little more context for all of this in those earlier posts.) The statistical inference in the study will be based on the estimated posterior distributions from a Bayesian model, so it seems like we’d like those distributions to be as informative as possible. We need to set the sample size large enough to reduce the dispersion of those distributions to a helpful level.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A Bayesian analysis of a factorial design focusing on effect size estimates</title>
      <link>https://www.rdatagen.net/post/2021-10-12-analyzing-a-factorial-design-with-a-bayesian-shrinkage-model/</link>
      <pubDate>Tue, 12 Oct 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-10-12-analyzing-a-factorial-design-with-a-bayesian-shrinkage-model/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-10-12-analyzing-a-factorial-design-with-a-bayesian-shrinkage-model/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;Factorial study designs present a number of analytic challenges, not least of which is how to best understand whether simultaneously applying multiple interventions is beneficial. &lt;a href=&#34;https://www.rdatagen.net/post/2021-09-28-analyzing-a-factorial-trial-with-a-bayesian-model/&#34; target=&#34;_blank&#34;&gt;Last time&lt;/a&gt; I presented a possible approach that focuses on estimating the variance of effect size estimates using a Bayesian model. The scenario I used there focused on a hypothetical study evaluating two interventions with four different levels each. This time around, I am considering a proposed study to reduce emergency department (ED) use for patients living with dementia that I am actually involved with. This study would have three different interventions, but only two levels for each (i.e., yes or no), for a total of 8 arms. In this case - the model I proposed previously does not seem like it would work well; the posterior distributions based on the variance-based model turn out to be bi-modal in shape, making it quite difficult to interpret the findings. So, I decided to turn the focus away from variance and emphasize the effect size estimates for each arm compared to control.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Analyzing a factorial design by focusing on the variance of effect sizes</title>
      <link>https://www.rdatagen.net/post/2021-09-28-analyzing-a-factorial-trial-with-a-bayesian-model/</link>
      <pubDate>Tue, 28 Sep 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-09-28-analyzing-a-factorial-trial-with-a-bayesian-model/</guid>
      <description>&lt;p&gt;Way back in 2018, long before the pandemic, I &lt;a href=&#34;https://www.rdatagen.net/post/testing-many-interventions-in-a-single-experiment/&#34; target=&#34;_blank&#34;&gt;described&lt;/a&gt; a soon-to-be implemented &lt;code&gt;simstudy&lt;/code&gt; function &lt;code&gt;genMultiFac&lt;/code&gt; that facilitates the generation of multi-factorial study data. I &lt;a href=&#34;https://www.rdatagen.net/post/so-how-efficient-are-multifactorial-experiments-part/&#34; target=&#34;_blank&#34;&gt;followed up&lt;/a&gt; that post with a description of how we can use these types of efficient designs to answer multiple questions in the context of a single study.&lt;/p&gt;&#xA;&lt;p&gt;Fast forward three years, and I am thinking about these designs again for a new grant application that proposes to study simultaneously three interventions aimed at reducing emergency department (ED) use for people living with dementia. The primary interest is to evaluate each intervention on its own terms, but also to assess whether any combinations seem to be particularly effective. While this will be a fairly large cluster randomized trial with about 80 EDs being randomized to one of the 8 possible combinations, I was concerned about our ability to estimate the interaction effects of multiple interventions with sufficient precision to draw useful conclusions, particularly if the combined effects of two or three interventions are less than additive. (That is, two interventions may be better than one, but not twice as good.)&lt;/p&gt;</description>
    </item>
    <item>
      <title>Drawing the wrong conclusion about subgroups: a comparison of Bayes and frequentist methods</title>
      <link>https://www.rdatagen.net/post/2021-09-14-drawing-the-wrong-conclusion-a-comparison-of-bayes-and-frequentist-methods/</link>
      <pubDate>Tue, 14 Sep 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-09-14-drawing-the-wrong-conclusion-a-comparison-of-bayes-and-frequentist-methods/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-09-14-drawing-the-wrong-conclusion-a-comparison-of-bayes-and-frequentist-methods/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;In the previous &lt;a href=&#34;https://www.rdatagen.net/post/2021-08-31-subgroup-analysis-using-a-bayesian-hierarchical-model/&#34; target=&#34;_blank&#34;&gt;post&lt;/a&gt;, I simulated data from a hypothetical RCT that had heterogeneous treatment effects across subgroups defined by three covariates. I presented two Bayesian models, a strongly &lt;em&gt;pooled&lt;/em&gt; model and an &lt;em&gt;unpooled&lt;/em&gt; version, that could be used to estimate all the subgroup effects in a single model. I compared the estimates to a set of linear regression models that were estimated for each subgroup separately.&lt;/p&gt;&#xA;&lt;p&gt;My goal in doing these comparisons is to see how often we might draw the wrong conclusion about subgroup effects when we conduct these types of analyses. In a typical frequentist framework, the probability of making a mistake is usually considerably greater than the 5% error rate that we allow ourselves, because conducting multiple tests gives us more chances to make a mistake. By using Bayesian hierarchical models that share information across subgroups and more reasonably measure uncertainty, I wanted to see if we can reduce the chances of drawing the wrong conclusions.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Subgroup analysis using a Bayesian hierarchical model</title>
      <link>https://www.rdatagen.net/post/2021-08-31-subgroup-analysis-using-a-bayesian-hierarchical-model/</link>
      <pubDate>Tue, 31 Aug 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-08-31-subgroup-analysis-using-a-bayesian-hierarchical-model/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-08-31-subgroup-analysis-using-a-bayesian-hierarchical-model/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;I’m part of a team that recently submitted the results of a randomized clinical trial for publication in a journal. The overall findings of the study were inconclusive, and we certainly didn’t try to hide that fact in our paper. Of course, the story was a bit more complicated, as the RCT was conducted during various phases of the COVID-19 pandemic; the context in which the therapeutic treatment was provided changed over time. In particular, other new treatments became standard of care along the way, resulting in apparent heterogeneous treatment effects for the therapy we were studying. It appears as if the treatment we were studying might have been effective only in one period when alternative treatments were not available. While we planned to evaluate the treatment effect over time, it was not our primary planned analysis, and the journal objected to the inclusion of the these secondary analyses.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Posterior probability checking with rvars: a quick follow-up</title>
      <link>https://www.rdatagen.net/post/2021-08-17-quick-follow-up-on-posterior-probability-checks-with-rvars/</link>
      <pubDate>Tue, 17 Aug 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-08-17-quick-follow-up-on-posterior-probability-checks-with-rvars/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-08-17-quick-follow-up-on-posterior-probability-checks-with-rvars/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;This is a relatively brief addendum to last week’s &lt;a href=&#34;https://www.rdatagen.net/post/2021-08-10-fitting-your-model-is-only-the-begining-bayesian-posterior-probability-checks/&#34;&gt;post&lt;/a&gt;, where I described how the &lt;code&gt;rvar&lt;/code&gt; datatype implemented in the &lt;code&gt;R&lt;/code&gt; package &lt;code&gt;posterior&lt;/code&gt; makes it quite easy to perform posterior probability checks to assess goodness of fit. In the initial post, I generated data from a linear model and estimated parameters for a linear regression model, and, unsurprisingly, the model fit the data quite well. When I introduced a quadratic term into the data generating process and fit the same linear model (without a quadratic term), equally unsurprising, the model wasn’t a great fit.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Fitting your model is only the beginning: Bayesian posterior probability checks with rvars</title>
      <link>https://www.rdatagen.net/post/2021-08-10-fitting-your-model-is-only-the-begining-bayesian-posterior-probability-checks/</link>
      <pubDate>Mon, 09 Aug 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-08-10-fitting-your-model-is-only-the-begining-bayesian-posterior-probability-checks/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-08-10-fitting-your-model-is-only-the-begining-bayesian-posterior-probability-checks/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;Say we’ve collected data and estimated parameters of a model that give structure to the data. An important question to ask is whether the model is a reasonable approximation of the true underlying data generating process. If we did a good job, we should be able to turn around and generate data from the model itself that looks similar to the data we started with. And if we didn’t do such a great job, the newly generated data will diverge from the original.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Estimating a risk difference (and confidence intervals) using logistic regression</title>
      <link>https://www.rdatagen.net/post/2021-06-15-estimating-a-risk-difference-using-logistic-regression/</link>
      <pubDate>Tue, 15 Jun 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-06-15-estimating-a-risk-difference-using-logistic-regression/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-06-15-estimating-a-risk-difference-using-logistic-regression/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;The &lt;em&gt;odds ratio&lt;/em&gt; (OR) – the effect size parameter estimated in logistic regression – is notoriously difficult to interpret. It is a ratio of two quantities (odds, under different conditions) that are themselves ratios of probabilities. I think it is pretty clear that a very large or small OR implies a strong treatment effect, but translating that effect into a clinical context can be challenging, particularly since ORs cannot be mapped to unique probabilities.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Sample size determination in the context of Bayesian analysis</title>
      <link>https://www.rdatagen.net/post/2021-06-01-bayesian-power-analysis/</link>
      <pubDate>Tue, 01 Jun 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-06-01-bayesian-power-analysis/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-06-01-bayesian-power-analysis/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;Given my recent involvement with the design of a somewhat complex &lt;a href=&#34;https://www.rdatagen.net/post/2021-01-19-should-we-continue-recruiting-patients-an-application-of-bayesian-predictive-probabilities/&#34; target=&#34;_blank&#34;&gt;trial&lt;/a&gt; centered around a Bayesian data analysis, I am appreciating more and more that Bayesian approaches are a very real option for clinical trial design. A key element of any study design is sample size. While some would argue that sample size considerations are not critical to the Bayesian design (since Bayesian inference is agnostic to any pre-specified sample size and is not really affected by how frequently you look at the data along the way), it might be a bit of a challenge to submit a grant without telling the potential funders how many subjects you plan on recruiting (since that could have a rather big effect on the level of resources - financial and time - required.)&lt;/p&gt;</description>
    </item>
    <item>
      <title>Generating random lists of names with errors to explore fuzzy word matching</title>
      <link>https://www.rdatagen.net/post/2021-04-13-generating-random-lists-of-names-with-errors-to-explore-fuzzy-word-matching/</link>
      <pubDate>Tue, 13 Apr 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-04-13-generating-random-lists-of-names-with-errors-to-explore-fuzzy-word-matching/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-04-13-generating-random-lists-of-names-with-errors-to-explore-fuzzy-word-matching/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;Health data systems are not always perfect, a point that was made quite obvious when a study I am involved with required a matched list of nursing home residents taken from one system with set results from PCR tests for COVID-19 drawn from another. Name spellings for the same person from the second list were not always consistent across different PCR tests, nor were they always consistent with the cohort we were interested in studying defined by the first list. My research associate, Yifan Xu, and I were asked to see what we could do to help out.&lt;/p&gt;</description>
    </item>
    <item>
      <title>The case of three MAR mechanisms: when is multiple imputation mandatory?</title>
      <link>https://www.rdatagen.net/post/2021-03-30-some-cases-where-imputing-missing-data-matters/</link>
      <pubDate>Tue, 30 Mar 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-03-30-some-cases-where-imputing-missing-data-matters/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-03-30-some-cases-where-imputing-missing-data-matters/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;I thought I’d written about this before, but I searched through my posts and I couldn’t find what I was looking for. If I am repeating myself, my apologies. I &lt;a href=&#34;https://www.rdatagen.net/post/musings-on-missing-data/&#34; target=&#34;_blank&#34;&gt;explored&lt;/a&gt; missing data two years ago, using directed acyclic graphs (DAGs) to help understand the various missing data mechanisms (MAR, MCAR, and MNAR). The DAGs provide insight into when it is appropriate to use observed data to get unbiased estimates of population quantities even though some of the observations are missing information.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Framework for power analysis using simulation</title>
      <link>https://www.rdatagen.net/post/2021-03-16-framework-for-power-analysis-using-simulation/</link>
      <pubDate>Tue, 16 Mar 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-03-16-framework-for-power-analysis-using-simulation/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-03-16-framework-for-power-analysis-using-simulation/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;The &lt;a href=&#34;https://kgoldfeld.github.io/simstudy/index.html&#34; target=&#34;_blank&#34;&gt;simstudy&lt;/a&gt; package started as a collection of functions I developed as I found myself repeating many of the same types of simulations for different projects. It was a way of organizing my work that I decided to share with others in case they wanted a routine way to generate data as well. &lt;code&gt;simstudy&lt;/code&gt; has expanded a bit from that, but replicability is still a key motivation.&lt;/p&gt;&#xA;&lt;p&gt;What I have here is another attempt to document and organize a process that I find myself doing quite often - repeated data generation and model fitting. Whether I am conducting a power analysis using simulation or exploring operating characteristics of different models, I take a pretty similar approach. I refer to this structure when I am starting a new project, so I thought it would be nice to have it easily accessible online - and that way others might be able to refer to it as well.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Randomization tests make fewer assumptions and seem pretty intuitive</title>
      <link>https://www.rdatagen.net/post/2021-03-02-randomization-tests/</link>
      <pubDate>Tue, 02 Mar 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-03-02-randomization-tests/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-03-02-randomization-tests/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;I’m preparing a lecture on simulation for a statistical modeling class, and I plan on describing a couple of cases where simulation is intrinsic to the analytic method rather than as a tool for exploration and planning. MCMC methods used for Bayesian estimation, bootstrapping, and randomization tests all come to mind.&lt;/p&gt;&#xA;&lt;p&gt;Randomization tests are particularly interesting as an approach to conducting hypothesis tests, because they allow us to avoid making unrealistic assumptions. I’ve written about this &lt;a href=&#34;https://www.rdatagen.net/post/permutation-test-for-a-covid-19-pilot-nursing-home-study/&#34; target=&#34;_blank&#34;&gt;before&lt;/a&gt; under the rubric of a permutation test. The example I use here is a little a different; truth be told, the real reason I’m sharing is that I came up with a nice little animation to illustrate a simple randomization process. So, even if I decide not to include it in the lecture, at least you’ve seen it.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Visualizing the treatment effect with an ordinal outcome</title>
      <link>https://www.rdatagen.net/post/2021-02-16-visualizing-the-treatment-effect-when-outcome-is-ordinal/</link>
      <pubDate>Tue, 16 Feb 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-02-16-visualizing-the-treatment-effect-when-outcome-is-ordinal/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-02-16-visualizing-the-treatment-effect-when-outcome-is-ordinal/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;If it’s true that many readers of a journal article focus on the abstract, figures and tables while skimming the rest, it is particularly important tell your story with a well conceived graphic or two. Along with a group of collaborators, I am trying to figure out the best way to represent an ordered categorical outcome from an RCT. In this case, there are a lot of categories, so the images can get confusing. I’m sharing a few of the possibilities that I’ve tried so far, including the code.&lt;/p&gt;</description>
    </item>
    <item>
      <title>How useful is it to show uncertainty in a plot comparing proportions?</title>
      <link>https://www.rdatagen.net/post/2021-02-02-uncertainty-in-a-plot-comparing-proportions/</link>
      <pubDate>Tue, 02 Feb 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-02-02-uncertainty-in-a-plot-comparing-proportions/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-02-02-uncertainty-in-a-plot-comparing-proportions/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;I recently created a simple plot for a paper describing a pilot study of an intervention targeting depression. This small study was largely conducted to assess the feasibility and acceptability of implementing an existing intervention in a new population. The primary outcome measure that was collected was the proportion of patients in each study arm who remained depressed following the intervention. The plot of the study results that we included in the paper looked something like this:&lt;/p&gt;</description>
    </item>
    <item>
      <title>Finding answers faster for COVID-19: an application of Bayesian predictive probabilities</title>
      <link>https://www.rdatagen.net/post/2021-01-19-should-we-continue-recruiting-patients-an-application-of-bayesian-predictive-probabilities/</link>
      <pubDate>Tue, 19 Jan 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-01-19-should-we-continue-recruiting-patients-an-application-of-bayesian-predictive-probabilities/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/post/2021-01-19-should-we-continue-recruiting-patients-an-application-of-bayesian-predictive-probabilities/index.en_files/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;As we evaluate therapies for COVID-19 to help improve outcomes during the pandemic, researchers need to be able to make recommendations as quickly as possible. There really is no time to lose. The Data &amp;amp; Safety Monitoring Board (DSMB) of &lt;a href=&#34;https://bit.ly/3qhY2f5&#34; target=&#34;_blank&#34;&gt;COMPILE&lt;/a&gt;, a prospective individual patient data meta-analysis, recognizes this. They are regularly monitoring the data to determine if there is a sufficiently strong signal to indicate effectiveness of convalescent plasma (CP) for hospitalized patients not on ventilation.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Coming soon: effortlessly generate ordinal data without assuming proportional odds</title>
      <link>https://www.rdatagen.net/post/2021-01-05-coming-soon-new-feature-to-easily-generate-cumulative-odds-without-proportionality-assumption/</link>
      <pubDate>Tue, 05 Jan 2021 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2021-01-05-coming-soon-new-feature-to-easily-generate-cumulative-odds-without-proportionality-assumption/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;I’m starting off 2021 with my 99th post ever to introduce a new feature that will be incorporated into &lt;code&gt;simstudy&lt;/code&gt; soon to make it a bit easier to generate ordinal data without requiring an assumption of proportional odds. I should wait until this feature has been incorporated into the development version, but I want to put it out there in case any one has any further suggestions. In any case, having this out in plain view will motivate me to get back to work on the package.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Constrained randomization to evaulate the vaccine rollout in nursing homes</title>
      <link>https://www.rdatagen.net/post/2020-12-22-constrained-randomization-to-evaulate-the-vaccine-rollout-in-nursing-homes/</link>
      <pubDate>Tue, 22 Dec 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/2020-12-22-constrained-randomization-to-evaulate-the-vaccine-rollout-in-nursing-homes/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;On an incredibly heartening note, two COVID-19 vaccines have been approved for use in the US and other countries around the world. More are possibly on the way. The big challenge, at least here in the United States, is to convince people that these vaccines are safe and effective; we need people to get vaccinated as soon as they are able to slow the spread of this disease. I for one will not hesitate for a moment to get a shot when I have the opportunity, though I don’t think biostatisticians are too high on the priority list.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A Bayesian implementation of a latent threshold model</title>
      <link>https://www.rdatagen.net/post/a-latent-threshold-model-to-estimate-treatment-effects/</link>
      <pubDate>Tue, 08 Dec 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/a-latent-threshold-model-to-estimate-treatment-effects/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;In the &lt;a href=&#34;https://www.rdatagen.net/post/a-latent-threshold-model/&#34; target=&#34;_blank&#34;&gt;previous post&lt;/a&gt;, I described a latent threshold model that might be helpful if we want to dichotomize a continuous predictor but we don’t know the appropriate cut-off point. This was motivated by a need to identify a threshold of antibody levels present in convalescent plasma that is currently being tested as a therapy for hospitalized patients with COVID in a number of RCTs, including those that are particpating in the ongoing &lt;a href=&#34;https://bit.ly/3lTTc4Q&#34; target=&#34;_blank&#34;&gt;COMPILE meta-analysis&lt;/a&gt;.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A latent threshold model to dichotomize a continuous predictor</title>
      <link>https://www.rdatagen.net/post/a-latent-threshold-model/</link>
      <pubDate>Tue, 24 Nov 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/a-latent-threshold-model/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;This is the context. In the convalescent plasma pooled individual patient level meta-analysis we are conducting as part of the &lt;a href=&#34;https://bit.ly/3nBxPXd&#34; target=&#34;_blank&#34;&gt;COMPILE&lt;/a&gt; study, there is great interest in understanding the impact of antibody levels on outcomes. (I’ve described various aspects of the analysis in previous posts, most recently &lt;a href=&#34;https://www.rdatagen.net/post/a-frequentist-bayesian-exploring-frequentist-properties-of-bayesian-models/&#34; target=&#34;_blank&#34;&gt;here&lt;/a&gt;). In other words, not all convalescent plasma is equal.&lt;/p&gt;&#xA;&lt;p&gt;If we had a clear measure of antibodies, we could model the relationship of these levels with the outcome of interest, such as health status as captured by the WHO 11-point scale or mortality, and call it a day. Unfortunately, at the moment, there is no single measure across the RCTs included in the meta-analysis (though that may change). Until now, the RCTs have used a range of measurement “platforms” (or technologies), which may measure different components of the convalescent plasma using different scales. Given these inconsistencies, it is challenging to build a straightforward model that simply estimates the relationship between antibody levels and clinical outcomes.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Exploring the properties of a Bayesian model using high performance computing</title>
      <link>https://www.rdatagen.net/post/a-frequentist-bayesian-exploring-frequentist-properties-of-bayesian-models/</link>
      <pubDate>Tue, 10 Nov 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/a-frequentist-bayesian-exploring-frequentist-properties-of-bayesian-models/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;An obvious downside to estimating Bayesian models is that it can take a considerable amount of time merely to fit a model. And if you need to estimate the same model repeatedly, that considerable amount becomes a prohibitive amount. In this post, which is part of a series (last one &lt;a href=&#34;https://bit.ly/31kCCDV&#34; target=&#34;_blank&#34;&gt;here&lt;/a&gt;) where I’ve been describing various aspects of the Bayesian analyses we plan to conduct for the &lt;a href=&#34;https://bit.ly/31hDwB0&#34; target=&#34;_blank&#34;&gt;COMPILE&lt;/a&gt; meta-analysis of convalescent plasma RCTs, I’ll present a somewhat elaborate model to illustrate how we have addressed these computing challenges to explore the properties of these models.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A refined brute force method to inform simulation of ordinal response data</title>
      <link>https://www.rdatagen.net/post/can-empirical-mean-and-variance-data-inform-simulation-of-ordinal-response-variables/</link>
      <pubDate>Tue, 27 Oct 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/can-empirical-mean-and-variance-data-inform-simulation-of-ordinal-response-variables/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;Francisco, a researcher from Spain, reached out to me with a challenge. He is interested in exploring various models that estimate correlation across multiple responses to survey questions. This is the context:&lt;/p&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;He doesn’t have access to actual data, so to explore analytic methods he needs to simulate responses.&lt;/li&gt;&#xA;&lt;li&gt;It would be ideal if the simulated data reflect the properties of real-world responses, some of which can be gleaned from the literature.&lt;/li&gt;&#xA;&lt;li&gt;The studies he’s found report only means and standard deviations of the ordinal data, along with the correlation matrices, &lt;em&gt;but not probability distributions of the responses&lt;/em&gt;.&lt;/li&gt;&#xA;&lt;li&gt;He’s considering &lt;code&gt;simstudy&lt;/code&gt; for his simulations, but the function &lt;code&gt;genOrdCat&lt;/code&gt; requires a set of probabilities for each response measure; it doesn’t seem like simstudy will be helpful here.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&lt;p&gt;Ultimately, we needed to figure out if we can we use the empirical means and standard deviations to derive probabilities that will yield those same means and standard deviations when the data are simulated. I thought about this for a bit, and came up with a bit of a work-around; the approach seems to work decently and doesn’t require any outrageous assumptions.&lt;/p&gt;</description>
    </item>
    <item>
      <title>simstudy just got a little more dynamic: version 0.2.1</title>
      <link>https://www.rdatagen.net/post/simstudy-just-got-a-little-more-dynamic-version-0-2-0/</link>
      <pubDate>Tue, 13 Oct 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/simstudy-just-got-a-little-more-dynamic-version-0-2-0/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;&lt;code&gt;simstudy&lt;/code&gt; version 0.2.1 has just been submitted to &lt;a href=&#34;https://cran.rstudio.com/web/packages/simstudy/&#34; target=&#34;_blank&#34;&gt;CRAN&lt;/a&gt;. Along with this release, the big news is that I’ve been joined by Jacob Wujciak-Jens as a co-author of the package. He initially reached out to me from Germany with some suggestions for improvements, we had a little back and forth, and now here we are. He has substantially reworked the underbelly of &lt;code&gt;simstudy&lt;/code&gt;, making the package much easier to maintain, and positioning it for much easier extension. And he implemented an entire system of formalized tests using &lt;a href=&#34;https://testthat.r-lib.org/&#34; target=&#34;_blank&#34;&gt;testthat&lt;/a&gt; and &lt;a href=&#34;https://cran.r-project.org/web/packages/hedgehog/vignettes/hedgehog.html&#34; target=&#34;_blank&#34;&gt;hedgehog&lt;/a&gt;; that was always my intention, but I never had the wherewithal to pull it off, and Jacob has done that. But, most importantly, it is much more fun to collaborate on this project than to toil away on my own.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Permuted block randomization using simstudy</title>
      <link>https://www.rdatagen.net/post/permuted-block-randomization-using-simstudy/</link>
      <pubDate>Tue, 29 Sep 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/permuted-block-randomization-using-simstudy/</guid>
      <description>&lt;p&gt;Along with preparing power analyses and statistical analysis plans (SAPs), generating study randomization lists is something a practicing biostatistician is occasionally asked to do. While not a particularly interesting activity, it offers the opportunity to tackle a small programming challenge. The title is a little misleading because you should probably skip all this and just use the &lt;code&gt;blockrand&lt;/code&gt; package if you want to generate randomization schemes; don’t try to reinvent the wheel. But, I can’t resist. Since I was recently asked to generate such a list, I’ve been wondering how hard it would be to accomplish this using &lt;code&gt;simstudy&lt;/code&gt;. There are already built-in functions for simulating stratified randomization schemes, so maybe it could be a good solution. The key element that is missing from simstudy, of course, is the permuted block setup.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Generating probabilities for ordinal categorical data</title>
      <link>https://www.rdatagen.net/post/generating-probabilities-for-ordinal-categorical-data/</link>
      <pubDate>Tue, 15 Sep 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/generating-probabilities-for-ordinal-categorical-data/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;Over the past couple of months, I’ve been describing various aspects of the simulations that we’ve been doing to get ready for a meta-analysis of convalescent plasma treatment for hospitalized patients with COVID-19, most recently &lt;a href=&#34;https://www.rdatagen.net/post/diagnosing-and-dealing-with-estimation-issues-in-the-bayesian-meta-analysis/&#34; target=&#34;_blank&#34;&gt;here&lt;/a&gt;. As I continue to do that, I want to provide motivation and code for a small but important part of the data generating process, which involves creating probabilities for ordinal categorical outcomes using a Dirichlet distribution.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Diagnosing and dealing with degenerate estimation in a Bayesian meta-analysis</title>
      <link>https://www.rdatagen.net/post/diagnosing-and-dealing-with-estimation-issues-in-the-bayesian-meta-analysis/</link>
      <pubDate>Tue, 01 Sep 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/diagnosing-and-dealing-with-estimation-issues-in-the-bayesian-meta-analysis/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;The federal government recently granted emergency approval for the use of antibody rich blood plasma when treating hospitalized COVID-19 patients. This announcement is &lt;a href=&#34;https://www.statnews.com/2020/08/24/trump-opened-floodgates-convalescent-plasma-too-soon/&#34; target=&#34;_blank&#34;&gt;unfortunate&lt;/a&gt;, because we really don’t know if this promising treatment works. The best way to determine this, of course, is to conduct an experiment, though this approval makes this more challenging to do; with the general availability of convalescent plasma (CP), there may be resistance from patients and providers against participating in a randomized trial. The emergency approval sends the incorrect message that the treatment is definitively effective. Why would a patient take the risk of receiving a placebo when they have almost guaranteed access to the therapy?&lt;/p&gt;</description>
    </item>
    <item>
      <title>Generating data from a truncated distribution</title>
      <link>https://www.rdatagen.net/post/generating-data-from-a-truncated-distribution/</link>
      <pubDate>Tue, 18 Aug 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/generating-data-from-a-truncated-distribution/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;A researcher reached out to me the other day to see if the &lt;code&gt;simstudy&lt;/code&gt; package provides a quick and easy way to generate data from a truncated distribution. Other than the &lt;code&gt;noZeroPoisson&lt;/code&gt; distribution option (which is a &lt;em&gt;very&lt;/em&gt; specific truncated distribution), there is no way to do this directly. You can always generate data from the full distribution and toss out the observations that fall outside of the truncation range, but this is not exactly efficient, and in practice can get a little messy. I’ve actually had it in the back of my mind to add something like this to &lt;code&gt;simstudy&lt;/code&gt;, but have hesitated because it might mean changing (or at least adding to) the &lt;code&gt;defData&lt;/code&gt; table structure.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A hurdle model for COVID-19 infections in nursing homes</title>
      <link>https://www.rdatagen.net/post/a-hurdle-model-for-covid-19-infections-in-nursing-homes-sample-size-considerations/</link>
      <pubDate>Tue, 04 Aug 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/a-hurdle-model-for-covid-19-infections-in-nursing-homes-sample-size-considerations/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;Late last &lt;a href=&#34;https://www.rdatagen.net/post/adding-mixture-distributions-to-simstudy/&#34; target=&#34;blank&#34;&gt;year&lt;/a&gt;, I added a &lt;em&gt;mixture&lt;/em&gt; distribution to the &lt;code&gt;simstudy&lt;/code&gt; package, largely motivated to accommodate &lt;em&gt;zero-inflated&lt;/em&gt; Poisson or negative binomial distributions. (I really thought I had added this two years ago - but time is moving so slowly these days.) These distributions are useful when modeling count data, but we anticipate observing more than the expected frequency of zeros that would arise from a non-inflated (i.e. “regular”) Poisson or negative binomial distribution.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A Bayesian model for a simulated meta-analysis</title>
      <link>https://www.rdatagen.net/post/a-bayesian-model-for-a-simulated-meta-analysis/</link>
      <pubDate>Tue, 21 Jul 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/a-bayesian-model-for-a-simulated-meta-analysis/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;This is essentially an addendum to the previous &lt;a href=&#34;https://www.rdatagen.net/post/simulating-mutliple-studies-to-simulate-a-meta-analysis/&#34; target=&#34;blank&#34;&gt;post&lt;/a&gt; where I simulated data from multiple RCTs to explore an analytic method to pool data across different studies. In that post, I used the &lt;code&gt;nlme&lt;/code&gt; package to conduct a meta-analysis based on individual level data of 12 studies. Here, I am presenting an alternative hierarchical modeling approach that uses the Bayesian package &lt;code&gt;rstan&lt;/code&gt;.&lt;/p&gt;&#xA;&lt;div id=&#34;create-the-data-set&#34; class=&#34;section level3&#34;&gt;&#xA;&lt;h3&gt;Create the data set&lt;/h3&gt;&#xA;&lt;p&gt;We’ll use the exact same data generating process as &lt;a href=&#34;https://www.rdatagen.net/post/simulating-mutliple-studies-to-simulate-a-meta-analysis/&#34; target=&#34;blank&#34;&gt;described&lt;/a&gt; in some detail in the previous post.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Simulating multiple RCTs to simulate a meta-analysis</title>
      <link>https://www.rdatagen.net/post/simulating-mutliple-studies-to-simulate-a-meta-analysis/</link>
      <pubDate>Tue, 07 Jul 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/simulating-mutliple-studies-to-simulate-a-meta-analysis/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;I am currently involved with an RCT that is struggling to recruit eligible patients (by no means an unusual problem), increasing the risk that findings might be inconclusive. A possible solution to this conundrum is to find similar, ongoing trials with the aim of pooling data in a single analysis, to conduct a &lt;em&gt;meta-analysis&lt;/em&gt; of sorts.&lt;/p&gt;&#xA;&lt;p&gt;In an ideal world, this theoretical collection of sites would have joined forces to develop a single study protocol, but often there is no structure or funding mechanism to make that happen. However, this group of studies may be similar enough - based on the target patient population, study inclusion and exclusion criteria, therapy protocols, comparison or control condition, randomization scheme, and outcome measurement - that it might be reasonable to estimate a single treatment effect and some measure of uncertainty.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Consider a permutation test for a small pilot study</title>
      <link>https://www.rdatagen.net/post/permutation-test-for-a-covid-19-pilot-nursing-home-study/</link>
      <pubDate>Tue, 23 Jun 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/permutation-test-for-a-covid-19-pilot-nursing-home-study/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;Recently I &lt;a href=&#34;https://www.rdatagen.net/post/what-can-we-really-expect-to-learn-from-a-pilot-study/&#34;&gt;wrote&lt;/a&gt; about the challenges of trying to learn too much from a small pilot study, even if it is a randomized controlled trial. There are limitations on how much you can learn about a treatment effect given the small sample size and relatively high variability of the estimate. However, the temptation for researchers is usually just too great; it is only natural to want to see if there is any kind of signal of an intervention effect, even though the pilot study is focused on questions of feasibility and acceptability.&lt;/p&gt;</description>
    </item>
    <item>
      <title>When proportional odds is a poor assumption, collapsing categories is probably not going to save you</title>
      <link>https://www.rdatagen.net/post/more-fun-with-ordinal-scales-combining-categories-may-not-make-solve-the-problem-of-non-proportionality/</link>
      <pubDate>Tue, 09 Jun 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/more-fun-with-ordinal-scales-combining-categories-may-not-make-solve-the-problem-of-non-proportionality/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;Continuing the discussion on cumulative odds models I started &lt;a href=&#34;https://www.rdatagen.net/post/the-advantage-of-increasing-the-number-of-categories-in-an-ordinal-outcome/&#34;&gt;last time&lt;/a&gt;, I want to investigate a solution I always assumed would help mitigate a failure to meet the proportional odds assumption. I’ve believed if there is a large number of categories and the relative cumulative odds between two groups don’t appear proportional across all categorical levels, then a reasonable approach is to reduce the number of categories. In other words, fewer categories translates to proportional odds. I’m not sure what led me to this conclusion, but in this post I’ve created some simulations that seem to throw cold water on that idea.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Considering the number of categories in an ordinal outcome</title>
      <link>https://www.rdatagen.net/post/the-advantage-of-increasing-the-number-of-categories-in-an-ordinal-outcome/</link>
      <pubDate>Tue, 26 May 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/the-advantage-of-increasing-the-number-of-categories-in-an-ordinal-outcome/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;In two Covid-19-related trials I’m involved with, the primary or key secondary outcome is the status of a patient at 14 days based on a World Health Organization ordered rating scale. In this particular ordinal scale, there are 11 categories ranging from 0 (uninfected) to 10 (death). In between, a patient can be infected but well enough to remain at home, hospitalized with milder symptoms, or hospitalized with severe disease. If the patient is hospitalized with severe disease, there are different stages of oxygen support the patient can be receiving, such as high flow oxygen or mechanical ventilation.&lt;/p&gt;</description>
    </item>
    <item>
      <title>To stratify or not? It might not actually matter...</title>
      <link>https://www.rdatagen.net/post/to-stratify-or-not-to-stratify/</link>
      <pubDate>Tue, 12 May 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/to-stratify-or-not-to-stratify/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;Continuing with the theme of &lt;em&gt;exploring small issues that come up in trial design&lt;/em&gt;, I recently used simulation to assess the impact of stratifying (or not) in the context of a multi-site Covid-19 trial with a binary outcome. The investigators are concerned that baseline health status will affect the probability of an outcome event, and are interested in randomizing by health status. The goal is to ensure balance across the two treatment arms with respect to this important variable. This randomization would be paired with an estimation model that adjusts for health status.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Simulation for power in designing cluster randomized trials</title>
      <link>https://www.rdatagen.net/post/simulation-for-power-calculations-in-designing-cluster-randomized-trials/</link>
      <pubDate>Tue, 28 Apr 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/simulation-for-power-calculations-in-designing-cluster-randomized-trials/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;As a biostatistician, I like to be involved in the design of a study as early as possible. I always like to say that I hope one of the first conversations an investigator has is with me, so that I can help clarify the research questions before getting into the design questions related to measurement, unit of randomization, and sample size. In the worst case scenario - and this actually doesn’t happen to me any more - a researcher would approach me after everything is done except the analysis. (I guess this is the appropriate time to pull out the quote made by the famous statistician Ronald Fisher: “To consult the statistician after an experiment is finished is often merely to ask him to conduct a post-mortem examination. He can perhaps say what the experiment died of.”)&lt;/p&gt;</description>
    </item>
    <item>
      <title>Yes, unbalanced randomization can improve power, in some situations</title>
      <link>https://www.rdatagen.net/post/unbalanced-randomization-can-improve-power-in-some-situations/</link>
      <pubDate>Tue, 14 Apr 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/unbalanced-randomization-can-improve-power-in-some-situations/</guid>
      <description>&lt;p&gt;Last time I provided some simulations that &lt;a href=&#34;https://www.rdatagen.net/post/can-unbalanced-randomization-improve-power/&#34;&gt;suggested&lt;/a&gt; that there might not be any efficiency-related benefits to using unbalanced randomization when the outcome is binary. This is a quick follow-up to provide a counter-example where the outcome in a two-group comparison is continuous. If the groups have different amounts of variability, intuitively it makes sense to allocate more patients to the more variable group. Doing this should reduce the variability in the estimate of the mean for that group, which in turn could improve the power of the test.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Can unbalanced randomization improve power?</title>
      <link>https://www.rdatagen.net/post/can-unbalanced-randomization-improve-power/</link>
      <pubDate>Tue, 31 Mar 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/can-unbalanced-randomization-improve-power/</guid>
      <description>&lt;p&gt;Of course, we’re all thinking about one thing these days, so it seems particularly inconsequential to be writing about anything that doesn’t contribute to solving or addressing in some meaningful way this pandemic crisis. But, I find that working provides a balm from reading and hearing all day about the events swirling around us, both here and afar. (I am in NYC, where things are definitely swirling.) And for me, working means blogging, at least for a few hours every couple of weeks.&lt;/p&gt;</description>
    </item>
    <item>
      <title>When you want more than a chi-squared test, consider a measure of association</title>
      <link>https://www.rdatagen.net/post/when-a-chi-squared-statistic-is-not-enough-a-measure-of-association-for-contingency-tables/</link>
      <pubDate>Tue, 17 Mar 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/when-a-chi-squared-statistic-is-not-enough-a-measure-of-association-for-contingency-tables/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;In my last &lt;a href=&#34;https://www.rdatagen.net/post/to-report-a-p-value-or-not-the-case-of-a-contingency-table/&#34;&gt;post&lt;/a&gt;, I made the point that p-values should not necessarily be considered sufficient evidence (or evidence at all) in drawing conclusions about associations we are interested in exploring. When it comes to contingency tables that represent the outcomes for two categorical variables, it isn’t so obvious what measure of association should augment (or replace) the &lt;span class=&#34;math inline&#34;&gt;\(\chi^2\)&lt;/span&gt; statistic.&lt;/p&gt;&#xA;&lt;p&gt;I described a model-based measure of effect to quantify the strength of an association in the particular case where one of the categorical variables is ordinal. This can arise, for example, when we want to compare Likert-type responses across multiple groups. The measure of effect I focused on - the cumulative proportional odds - is quite useful, but is potentially limited for two reasons. First, the proportional odds assumption may not be reasonable, potentially leading to biased estimates. Second, both factors may be nominal (i.e. not ordinal), it which case cumulative odds model is inappropriate.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Alternatives to reporting a p-value: the case of a contingency table</title>
      <link>https://www.rdatagen.net/post/to-report-a-p-value-or-not-the-case-of-a-contingency-table/</link>
      <pubDate>Tue, 03 Mar 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/to-report-a-p-value-or-not-the-case-of-a-contingency-table/</guid>
      <description>&lt;p&gt;I frequently find myself in discussions with collaborators about the merits of reporting p-values, particularly in the context of pilot studies or exploratory analysis. Over the past several years, the &lt;a href=&#34;https://www.amstat.org/&#34;&gt;&lt;em&gt;American Statistical Association&lt;/em&gt;&lt;/a&gt; has made several strong statements about the need to consider approaches that measure the strength of evidence or uncertainty that don’t necessarily rely on p-values. In &lt;a href=&#34;https://amstat.tandfonline.com/doi/full/10.1080/00031305.2016.1154108&#34;&gt;2016&lt;/a&gt;, the ASA attempted to clarify the proper use and interpretation of the p-value by highlighting key principles “that could improve the conduct or interpretation of quantitative science, according to widespread consensus in the statistical community.” These principles are worth noting here in case you don’t make it over to the original paper:&lt;/p&gt;</description>
    </item>
    <item>
      <title>Clustered randomized trials and the design effect</title>
      <link>https://www.rdatagen.net/post/what-exactly-is-the-design-effect/</link>
      <pubDate>Tue, 18 Feb 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/what-exactly-is-the-design-effect/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;I am always saying that simulation can help illuminate interesting statistical concepts or ideas. Such an exploration might provide some insight into the concept of the &lt;em&gt;design effect&lt;/em&gt;, which underlies clustered randomized trial designs. I’ve written about clustered-related methods so much on this blog that I won’t provide links - just peruse the list of entries on the home page and you are sure to spot a few. But, I haven’t written explicitly about the design effect.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Analysing an open cohort stepped-wedge clustered trial with repeated individual binary outcomes</title>
      <link>https://www.rdatagen.net/post/analyzing-the-open-cohort-stepped-wedge-trial-with-binary-outcomes/</link>
      <pubDate>Tue, 04 Feb 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/analyzing-the-open-cohort-stepped-wedge-trial-with-binary-outcomes/</guid>
      <description>&lt;p&gt;I am currently wrestling with how to analyze data from a stepped-wedge designed cluster randomized trial. A few factors make this analysis particularly interesting. First, we want to allow for the possibility that between-period site-level correlation will decrease (or decay) over time. Second, there is possibly additional clustering at the patient level since individual outcomes will be measured repeatedly over time. And third, given that these outcomes are binary, there are no obvious software tools that can handle generalized linear models with this particular variance structure we want to model. (If I have missed something obvious with respect to modeling options, please let me know.)&lt;/p&gt;</description>
    </item>
    <item>
      <title>A brief account (via simulation) of the ROC (and its AUC)</title>
      <link>https://www.rdatagen.net/post/a-simple-explanation-of-what-the-roc-and-auc-represent/</link>
      <pubDate>Tue, 21 Jan 2020 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/a-simple-explanation-of-what-the-roc-and-auc-represent/</guid>
      <description>&lt;p&gt;The ROC (receiver operating characteristic) curve visually depicts the ability of a measure or classification model to distinguish two groups. The area under the ROC (AUC), quantifies the extent of that ability. My goal here is to describe as simply as possible a process that serves as a foundation for the ROC, and to provide an interpretation of the AUC that is defined by that curve.&lt;/p&gt;&#xA;&lt;div id=&#34;a-prediction-problem&#34; class=&#34;section level2&#34;&gt;&#xA;&lt;h2&gt;A prediction problem&lt;/h2&gt;&#xA;&lt;p&gt;The classic application for the ROC is a medical test designed to identify individuals with a particular medical condition or disease. The population is comprised of two groups of individuals: those with the condition and those without. What we want is some sort of diagnostic tool (such as a blood test or diagnostic scan) that will identify which group a particular patient belongs to. The question is how well does that tool or measure help us distinguish between the two groups? The ROC (and AUC) is designed to help answer that question.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Repeated measures can improve estimation when we only care about a single endpoint</title>
      <link>https://www.rdatagen.net/post/using-repeated-measures-might-improve-effect-estimation-even-when-single-endpoint-is-the-focus/</link>
      <pubDate>Tue, 10 Dec 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/using-repeated-measures-might-improve-effect-estimation-even-when-single-endpoint-is-the-focus/</guid>
      <description>&lt;p&gt;I’m participating in the design of a new study that will evaluate interventions aimed at reducing both pain and opioid use for patients on dialysis. This study is likely to be somewhat complicated, possibly involving multiple clusters, multiple interventions, a sequential and/or adaptive randomization scheme, and a composite binary outcome. I’m not going into any of that here.&lt;/p&gt;&#xA;&lt;p&gt;There &lt;em&gt;is&lt;/em&gt; one issue that should be fairly generalizable to other studies. It is likely that individual measures will be collected repeatedly over time but the primary outcome of interest will be the measure collected during the last follow-up period. I wanted to explore what, if anything, can be gained by analyzing all of the available data rather than focusing only the final end point.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Adding a &#34;mixture&#34; distribution to the simstudy package</title>
      <link>https://www.rdatagen.net/post/adding-mixture-distributions-to-simstudy/</link>
      <pubDate>Tue, 26 Nov 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/adding-mixture-distributions-to-simstudy/</guid>
      <description>&lt;p&gt;I am contemplating adding a new distribution option to the package &lt;code&gt;simstudy&lt;/code&gt; that would allow users to define a new variable as a mixture of previously defined (or already generated) variables. I think the easiest way to explain how to apply the new &lt;em&gt;mixture&lt;/em&gt; option is to step through a few examples and see it in action.&lt;/p&gt;&#xA;&lt;div id=&#34;specifying-the-mixture-distribution&#34; class=&#34;section level3&#34;&gt;&#xA;&lt;h3&gt;Specifying the “mixture” distribution&lt;/h3&gt;&#xA;&lt;p&gt;As defined here, a mixture of variables is a random draw from a set of variables based on a defined set of probabilities. For example, if we have two variables, &lt;span class=&#34;math inline&#34;&gt;\(x_1\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(x_2\)&lt;/span&gt;, we have a mixture if, for any particular observation, we take &lt;span class=&#34;math inline&#34;&gt;\(x_1\)&lt;/span&gt; with probability &lt;span class=&#34;math inline&#34;&gt;\(p_1\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(x_2\)&lt;/span&gt; with probability &lt;span class=&#34;math inline&#34;&gt;\(p_2\)&lt;/span&gt;, where &lt;span class=&#34;math inline&#34;&gt;\(\sum_i{p_i} = 1\)&lt;/span&gt;, &lt;span class=&#34;math inline&#34;&gt;\(i \in (1, 2)\)&lt;/span&gt;. So, if we have already defined &lt;span class=&#34;math inline&#34;&gt;\(x_1\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(x_2\)&lt;/span&gt; using the &lt;code&gt;defData&lt;/code&gt; function, we can create a third variable &lt;span class=&#34;math inline&#34;&gt;\(x_{mix}\)&lt;/span&gt; with this definition:&lt;/p&gt;</description>
    </item>
    <item>
      <title>What can we really expect to learn from a pilot study?</title>
      <link>https://www.rdatagen.net/post/what-can-we-really-expect-to-learn-from-a-pilot-study/</link>
      <pubDate>Tue, 12 Nov 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/what-can-we-really-expect-to-learn-from-a-pilot-study/</guid>
      <description>&lt;p&gt;I am involved with a very interesting project - the &lt;a href=&#34;https://impactcollaboratory.org/&#34;&gt;NIA IMPACT Collaboratory&lt;/a&gt; - where a primary goal is to fund a large group of pragmatic pilot studies to investigate promising interventions to improve health care and quality of life for people living with Alzheimer’s disease and related dementias. One of my roles on the project team is to advise potential applicants on the development of their proposals. In order to provide helpful advice, it is important that we understand what we should actually expect to learn from a relatively small pilot study of a new intervention.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Any one interested in a function to quickly generate data with many predictors?</title>
      <link>https://www.rdatagen.net/post/any-one-interested-in-a-function-to-quickly-generate-data-with-many-predictors/</link>
      <pubDate>Tue, 29 Oct 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/any-one-interested-in-a-function-to-quickly-generate-data-with-many-predictors/</guid>
      <description>&lt;p&gt;A couple of months ago, I was contacted about the possibility of creating a simple function in &lt;code&gt;simstudy&lt;/code&gt; to generate a large dataset that could include possibly 10’s or 100’s of potential predictors and an outcome. In this function, only a subset of the variables would actually be predictors. The idea is to be able to easily generate data for exploring ridge regression, Lasso regression, or other “regularization” methods. Alternatively, this can be used to very quickly generate correlated data (with one line of code) without going through the definition process.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Selection bias, death, and dying</title>
      <link>https://www.rdatagen.net/post/selection-bias-death-and-dying/</link>
      <pubDate>Tue, 15 Oct 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/selection-bias-death-and-dying/</guid>
      <description>&lt;p&gt;I am collaborating with a number of folks who think a lot about palliative or supportive care for people who are facing end-stage disease, such as advanced dementia, cancer, COPD, or congestive heart failure. A major concern for this population (which really includes just about everyone at some point) is the quality of life at the end of life and what kind of experiences, including interactions with the health care system, they have (and don’t have) before death.&lt;/p&gt;</description>
    </item>
    <item>
      <title>There&#39;s always at least two ways to do the same thing: an example generating 3-level hierarchical data using simstudy</title>
      <link>https://www.rdatagen.net/post/in-simstudy-as-in-r-there-s-always-at-least-two-ways-to-do-the-same-thing/</link>
      <pubDate>Thu, 03 Oct 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/in-simstudy-as-in-r-there-s-always-at-least-two-ways-to-do-the-same-thing/</guid>
      <description>&lt;p&gt;“I am working on a simulation study that requires me to generate data for individuals within clusters, but each individual will have repeated measures (say baseline and two follow-ups). I’m new to simstudy and have been going through the examples in R this afternoon, but I wondered if this was possible in the package, and if so whether you could offer any tips to get me started with how I would do this?”&lt;/p&gt;</description>
    </item>
    <item>
      <title>Simulating an open cohort stepped-wedge trial</title>
      <link>https://www.rdatagen.net/post/simulating-an-open-cohort-stepped-wedge-trial/</link>
      <pubDate>Tue, 17 Sep 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/simulating-an-open-cohort-stepped-wedge-trial/</guid>
      <description>&lt;p&gt;In a current multi-site study, we are using a stepped-wedge design to evaluate whether improved training and protocols can reduce prescriptions of anti-psychotic medication for home hospice care patients with advanced dementia. The study is officially called the Hospice Advanced Dementia Symptom Management and Quality of Life (HAS-QOL) Stepped Wedge Trial. Unlike my previous work with &lt;a href=&#34;https://www.rdatagen.net/post/alternatives-to-stepped-wedge-designs/&#34;&gt;stepped-wedge designs&lt;/a&gt;, where individuals were measured once in the course of the study, this study will collect patient outcomes from the home hospice care EHRs over time. This means that for some patients, the data collection period straddles the transition from control to intervention.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Analyzing a binary outcome arising out of within-cluster, pair-matched randomization</title>
      <link>https://www.rdatagen.net/post/analyzing-a-binary-outcome-in-a-study-with-within-cluster-pair-matched-randomization/</link>
      <pubDate>Tue, 03 Sep 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/analyzing-a-binary-outcome-in-a-study-with-within-cluster-pair-matched-randomization/</guid>
      <description>&lt;p&gt;A key motivating factor for the &lt;code&gt;simstudy&lt;/code&gt; package and much of this blog is that simulation can be super helpful in understanding how best to approach an unusual, or least unfamiliar, analytic problem. About six months ago, I &lt;a href=&#34;https://www.rdatagen.net/post/a-case-where-prospecitve-matching-may-limit-bias/&#34;&gt;described&lt;/a&gt; the DREAM Initiative (Diabetes Research, Education, and Action for Minorities), a study that used a slightly innovative randomization scheme to ensure that two comparison groups were evenly balanced across important covariates. At the time, we hadn’t finalized the analytic plan. But, now that we have started actually randomizing and recruiting (yes, in that order, oddly enough), it is important that we do that, with the help of a little simulation.&lt;/p&gt;</description>
    </item>
    <item>
      <title>simstudy updated to version 0.1.14: implementing Markov chains</title>
      <link>https://www.rdatagen.net/post/simstudy-1-14-update/</link>
      <pubDate>Tue, 20 Aug 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/simstudy-1-14-update/</guid>
      <description>&lt;p&gt;I’m developing study simulations that require me to generate a sequence of health status for a collection of individuals. In these simulations, individuals gradually grow sicker over time, though sometimes they recover slightly. To facilitate this, I am using a stochastic Markov process, where the probability of a health status at a particular time depends only on the previous health status (in the immediate past). While there are packages to do this sort of thing (see for example the &lt;a href=&#34;https://cran.r-project.org/web/packages/markovchain/index.html&#34;&gt;markovchain&lt;/a&gt; package), I hadn’t yet stumbled upon them while I was tackling my problem. So, I wrote my own functions, which I’ve now incorporated into the latest version of &lt;code&gt;simstudy&lt;/code&gt; that is now available on &lt;a href=&#34;https://cran.r-project.org/web/packages/simstudy/index.html&#34;&gt;CRAN&lt;/a&gt;. As a way of announcing the new release, here is a brief overview of Markov chains and the new functions. (See &lt;a href=&#34;https://cran.r-project.org/web/packages/simstudy/news/news.html&#34;&gt;here&lt;/a&gt; for a more complete list of changes.)&lt;/p&gt;</description>
    </item>
    <item>
      <title>Bayes models for estimation in stepped-wedge trials with non-trivial ICC patterns</title>
      <link>https://www.rdatagen.net/post/bayes-model-to-estimate-stepped-wedge-trial-with-non-trivial-icc-structure/</link>
      <pubDate>Tue, 06 Aug 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/bayes-model-to-estimate-stepped-wedge-trial-with-non-trivial-icc-structure/</guid>
      <description>&lt;p&gt;Continuing a series of posts discussing the structure of intra-cluster correlations (ICC’s) in the context of a stepped-wedge trial, this latest edition is primarily interested in fitting Bayesian hierarchical models for more complex cases (though I do talk a bit more about the linear mixed effects models). The first two posts in the series focused on generating data to simulate various scenarios; the &lt;a href=&#34;https://www.rdatagen.net/post/estimating-treatment-effects-and-iccs-for-stepped-wedge-designs/&#34;&gt;third post&lt;/a&gt; considered linear mixed effects and Bayesian hierarchical models to estimate ICC’s under the simplest scenario of constant between-period ICC’s. Throughout this post, I use code drawn from the previous one; I am not repeating much of it here for brevity’s sake. So, if this is all new, it is probably worth &lt;a href=&#34;https://www.rdatagen.net/post/estimating-treatment-effects-and-iccs-for-stepped-wedge-designs/&#34;&gt;glancing at&lt;/a&gt; before continuing on.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Estimating treatment effects (and ICCs) for stepped-wedge designs</title>
      <link>https://www.rdatagen.net/post/estimating-treatment-effects-and-iccs-for-stepped-wedge-designs/</link>
      <pubDate>Tue, 16 Jul 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/estimating-treatment-effects-and-iccs-for-stepped-wedge-designs/</guid>
      <description>&lt;p&gt;In the last two posts, I introduced the notion of time-varying intra-cluster correlations in the context of stepped-wedge study designs. (See &lt;a href=&#34;https://www.rdatagen.net/post/intra-cluster-correlations-over-time/&#34;&gt;here&lt;/a&gt; and &lt;a href=&#34;https://www.rdatagen.net/post/varying-intra-cluster-correlations-over-time/&#34;&gt;here&lt;/a&gt;). Though I generated lots of data for those posts, I didn’t fit any models to see if I could recover the estimates and any underlying assumptions. That’s what I am doing now.&lt;/p&gt;&#xA;&lt;p&gt;My focus here is on the simplest case, where the ICC’s are constant over time and between time. Typically, I would just use a mixed-effects model to estimate the treatment effect and account for variability across clusters, which is easily done in &lt;code&gt;R&lt;/code&gt; using the &lt;code&gt;lme4&lt;/code&gt; package; if the outcome is continuous the function &lt;code&gt;lmer&lt;/code&gt; is appropriate. I thought, however, it would also be interesting to use the &lt;code&gt;rstan&lt;/code&gt; package to fit a Bayesian hierarchical model.&lt;/p&gt;</description>
    </item>
    <item>
      <title>More on those stepped-wedge design assumptions: varying intra-cluster correlations over time</title>
      <link>https://www.rdatagen.net/post/varying-intra-cluster-correlations-over-time/</link>
      <pubDate>Tue, 09 Jul 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/varying-intra-cluster-correlations-over-time/</guid>
      <description>&lt;p&gt;In my last &lt;a href=&#34;https://www.rdatagen.net/post/intra-cluster-correlations-over-time/&#34;&gt;post&lt;/a&gt;, I wrote about &lt;em&gt;within-&lt;/em&gt; and &lt;em&gt;between-period&lt;/em&gt; intra-cluster correlations in the context of stepped-wedge cluster randomized study designs. These are quite important to understand when figuring out sample size requirements (and models for analysis, which I’ll be writing about soon.) Here, I’m extending the constant ICC assumption I presented last time around by introducing some complexity into the correlation structure. Much of the code I am using can be found in last week’s post, so if anything seems a little unclear, hop over &lt;a href=&#34;https://www.rdatagen.net/post/intra-cluster-correlations-over-time/&#34;&gt;here&lt;/a&gt;.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Planning a stepped-wedge trial? Make sure you know what you&#39;re assuming about intra-cluster correlations ...</title>
      <link>https://www.rdatagen.net/post/intra-cluster-correlations-over-time/</link>
      <pubDate>Tue, 25 Jun 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/intra-cluster-correlations-over-time/</guid>
      <description>&lt;p&gt;A few weeks ago, I was at the annual meeting of the &lt;a href=&#34;https://rethinkingclinicaltrials.org/&#34;&gt;NIH Collaboratory&lt;/a&gt;, which is an innovative collection of collaboratory cores, demonstration projects, and NIH Institutes and Centers that is developing new models for implementing and supporting large-scale health services research. A study I am involved with - &lt;em&gt;Primary Palliative Care for Emergency Medicine&lt;/em&gt; - is one of the demonstration projects in this collaboratory.&lt;/p&gt;&#xA;&lt;p&gt;The second day of this meeting included four panels devoted to the design and analysis of embedded pragmatic clinical trials, and focused on the challenges of conducting rigorous research in the real-world context of a health delivery system. The keynote address that started off the day was presented by David Murray of NIH, who talked about the challenges and limitations of cluster randomized trials. (I’ve written before on issues related to clustered randomized trials, including &lt;a href=&#34;https://www.rdatagen.net/post/what-matters-more-in-a-cluster-randomized-trial-number-or-size/&#34;&gt;here&lt;/a&gt;.)&lt;/p&gt;</description>
    </item>
    <item>
      <title>Don&#39;t get too excited - it might just be regression to the mean</title>
      <link>https://www.rdatagen.net/post/regression-to-the-mean/</link>
      <pubDate>Tue, 11 Jun 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/regression-to-the-mean/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;It is always exciting to find an interesting pattern in the data that seems to point to some important difference or relationship. A while ago, one of my colleagues shared a figure with me that looked something like this:&lt;/p&gt;&#xA;&lt;p&gt;&lt;img src=&#34;https://www.rdatagen.net/post/2019-06-11-regression-to-the-mean.en_files/figure-html/unnamed-chunk-2-1.png&#34; width=&#34;672&#34; /&gt;&lt;/p&gt;&#xA;&lt;p&gt;It looks like something is going on. On average low scorers in the first period increased a bit in the second period, and high scorers decreased a bit. Something &lt;strong&gt;is&lt;/strong&gt; going on, but nothing specific to the data in question; it is just probability working its magic.&lt;/p&gt;</description>
    </item>
    <item>
      <title>simstudy update - stepped-wedge design treatment assignment</title>
      <link>https://www.rdatagen.net/post/simstudy-update-stepped-wedge-treatment-assignment/</link>
      <pubDate>Tue, 28 May 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/simstudy-update-stepped-wedge-treatment-assignment/</guid>
      <description>&lt;p&gt;&lt;code&gt;simstudy&lt;/code&gt; has just been updated (version 0.1.13 on &lt;a href=&#34;https://cran.rstudio.com/web/packages/simstudy/&#34;&gt;CRAN&lt;/a&gt;), and includes one interesting addition (and a couple of bug fixes). I am working on a post (or two) about intra-cluster correlations (ICCs) and stepped-wedge study designs (which I’ve written about &lt;a href=&#34;https://www.rdatagen.net/post/alternatives-to-stepped-wedge-designs/&#34;&gt;before&lt;/a&gt;), and I was getting tired of going through the convoluted process of generating data from a time-dependent treatment assignment process. So, I wrote a new function, &lt;code&gt;trtStepWedge&lt;/code&gt;, that should simplify things.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Generating and modeling over-dispersed binomial data</title>
      <link>https://www.rdatagen.net/post/overdispersed-binomial-data/</link>
      <pubDate>Tue, 14 May 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/overdispersed-binomial-data/</guid>
      <description>&lt;p&gt;A couple of weeks ago, I was inspired by a study to &lt;a href=&#34;https://www.rdatagen.net/post/what-matters-more-in-a-cluster-randomized-trial-number-or-size/&#34;&gt;write&lt;/a&gt; about a classic design issue that arises in cluster randomized trials: should we focus on the number of clusters or the size of those clusters? This trial, which is concerned with preventing opioid use disorder for at-risk patients in primary care clinics, has also motivated this second post, which concerns another important issue - over-dispersion.&lt;/p&gt;&#xA;&lt;div id=&#34;a-count-outcome&#34; class=&#34;section level3&#34;&gt;&#xA;&lt;h3&gt;A count outcome&lt;/h3&gt;&#xA;&lt;p&gt;In this study, one of the primary outcomes is the number of days of opioid use over a six-month follow-up period (to be recorded monthly by patient-report and aggregated for the six-month measure). While one might get away with assuming that the outcome is continuous, it really is not; it is a &lt;em&gt;count&lt;/em&gt; outcome, and the possible range is 0 to 180. There are two related questions here - what model will be used to analyze the data once the study is complete? And, how should we generate simulated data to estimate the power of the study?&lt;/p&gt;</description>
    </item>
    <item>
      <title>What matters more in a cluster randomized trial: number or size?</title>
      <link>https://www.rdatagen.net/post/what-matters-more-in-a-cluster-randomized-trial-number-or-size/</link>
      <pubDate>Tue, 30 Apr 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/what-matters-more-in-a-cluster-randomized-trial-number-or-size/</guid>
      <description>&lt;p&gt;I am involved with a trial of an intervention designed to prevent full-blown opioid use disorder for patients who may have an incipient opioid use problem. Given the nature of the intervention, it was clear the only feasible way to conduct this particular study is to randomize at the physician rather than the patient level.&lt;/p&gt;&#xA;&lt;p&gt;There was a concern that the number of patients eligible for the study might be limited, so that each physician might only have a handful of patients able to participate, if that many. A question arose as to whether we can make up for this limitation by increasing the number of physicians who participate? That is, what is the trade-off between number of clusters and cluster size?&lt;/p&gt;</description>
    </item>
    <item>
      <title>Musings on missing data</title>
      <link>https://www.rdatagen.net/post/musings-on-missing-data/</link>
      <pubDate>Tue, 02 Apr 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/musings-on-missing-data/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;I’ve been meaning to share an analysis I recently did to estimate the strength of the relationship between a young child’s ability to recognize emotions in others (e.g. teachers and fellow students) and her longer term academic success. The study itself is quite interesting (hopefully it will be published sometime soon), but I really wanted to write about it here as it involved the challenging problem of missing data in the context of heterogeneous effects (different across sub-groups) and clustering (by schools).&lt;/p&gt;</description>
    </item>
    <item>
      <title>A case where prospective matching may limit bias in a randomized trial</title>
      <link>https://www.rdatagen.net/post/a-case-where-prospecitve-matching-may-limit-bias/</link>
      <pubDate>Tue, 12 Mar 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/a-case-where-prospecitve-matching-may-limit-bias/</guid>
      <description>&lt;p&gt;Analysis is important, but study design is paramount. I am involved with the Diabetes Research, Education, and Action for Minorities (DREAM) Initiative, which is, among other things, estimating the effect of a group-based therapy program on weight loss for patients who have been identified as pre-diabetic (which means they have elevated HbA1c levels). The original plan was to randomize patients at a clinic to treatment or control, and then follow up with those assigned to the treatment group to see if they wanted to participate. The primary outcome is going to be measured using medical records, so those randomized to control (which basically means nothing special happens to them) will not need to interact with the researchers in any way.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A example in causal inference designed to frustrate: an estimate pretty much guaranteed to be biased</title>
      <link>https://www.rdatagen.net/post/dags-colliders-and-an-example-of-variance-bias-tradeoff/</link>
      <pubDate>Tue, 26 Feb 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/dags-colliders-and-an-example-of-variance-bias-tradeoff/</guid>
      <description>&lt;p&gt;I am putting together a brief lecture introducing causal inference for graduate students studying biostatistics. As part of this lecture, I thought it would be helpful to spend a little time describing directed acyclic graphs (DAGs), since they are an extremely helpful tool for communicating assumptions about the causal relationships underlying a researcher’s data.&lt;/p&gt;&#xA;&lt;p&gt;The strength of DAGs is that they help us think how these underlying relationships in the data might lead to biases in causal effect estimation, and suggest ways to estimate causal effects that eliminate these biases. (For a real introduction to DAGs, you could take a look at this &lt;a href=&#34;http://ftp.cs.ucla.edu/pub/stat_ser/r251.pdf&#34;&gt;paper&lt;/a&gt; by &lt;em&gt;Greenland&lt;/em&gt;, &lt;em&gt;Pearl&lt;/em&gt;, and &lt;em&gt;Robins&lt;/em&gt; or better yet take a look at Part I of this &lt;a href=&#34;https://www.hsph.harvard.edu/miguel-hernan/causal-inference-book/2015/&#34;&gt;book&lt;/a&gt; on causal inference by &lt;em&gt;Hernán&lt;/em&gt; and &lt;em&gt;Robins&lt;/em&gt;.)&lt;/p&gt;</description>
    </item>
    <item>
      <title>Using the uniform sum distribution to introduce probability</title>
      <link>https://www.rdatagen.net/post/a-fun-example-to-explore-probability/</link>
      <pubDate>Tue, 05 Feb 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/a-fun-example-to-explore-probability/</guid>
      <description>&lt;p&gt;I’ve never taught an intro probability/statistics course. If I ever did, I would certainly want to bring the underlying wonder of the subject to life. I’ve always found it almost magical the way mathematical formulation can be mirrored by computer simulation, the way proof can be guided by observed data generation processes, and the way DGPs can confirm analytic solutions.&lt;/p&gt;&#xA;&lt;p&gt;I would like to begin such a course with a somewhat unusual but accessible problem that would evoke these themes from the start. The concepts would not necessarily be immediately comprehensible, but rather would pique the interest of the students.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Correlated longitudinal data with varying time intervals</title>
      <link>https://www.rdatagen.net/post/correlated-longitudinal-data-with-varying-time-intervals/</link>
      <pubDate>Tue, 22 Jan 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/correlated-longitudinal-data-with-varying-time-intervals/</guid>
      <description>&lt;p&gt;I was recently contacted to see if &lt;code&gt;simstudy&lt;/code&gt; can create a data set of correlated outcomes that are measured over time, but at different intervals for each individual. The quick answer is there is no specific function to do this. However, if you are willing to assume an “exchangeable” correlation structure, where measurements far apart in time are just as correlated as measurements taken close together, then you could just generate individual-level random effects (intercepts and/or slopes) and pretty much call it a day. Unfortunately, the researcher had something more challenging in mind: he wanted to generate auto-regressive correlation, so that proximal measurements are more strongly correlated than distal measurements.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Considering sensitivity to unmeasured confounding: part 2</title>
      <link>https://www.rdatagen.net/post/what-does-it-mean-if-findings-are-sensitive-to-unmeasured-confounding-ii/</link>
      <pubDate>Thu, 10 Jan 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/what-does-it-mean-if-findings-are-sensitive-to-unmeasured-confounding-ii/</guid>
      <description>&lt;p&gt;In &lt;a href=&#34;https://www.rdatagen.net/post/what-does-it-mean-if-findings-are-sensitive-to-unmeasured-confounding/&#34;&gt;part 1&lt;/a&gt; of this 2-part series, I introduced the notion of &lt;em&gt;sensitivity to unmeasured confounding&lt;/em&gt; in the context of an observational data analysis. I argued that an estimate of an association between an observed exposure &lt;span class=&#34;math inline&#34;&gt;\(D\)&lt;/span&gt; and outcome &lt;span class=&#34;math inline&#34;&gt;\(Y\)&lt;/span&gt; is sensitive to unmeasured confounding if we can conceive of a reasonable alternative data generating process (DGP) that includes some unmeasured confounder that will generate the same observed distribution the observed data. I further argued that reasonableness can be quantified or parameterized by the two correlation coefficients &lt;span class=&#34;math inline&#34;&gt;\(\rho_{UD}\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(\rho_{UY}\)&lt;/span&gt;, which measure the strength of the relationship of the unmeasured confounder &lt;span class=&#34;math inline&#34;&gt;\(U\)&lt;/span&gt; with each of the observed measures. Alternative DGPs that are characterized by high correlation coefficients can be viewed as less realistic, and the observed data could be considered less sensitive to unmeasured confounding. On the other hand, DGPs characterized by lower correlation coefficients would be considered more sensitive.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Considering sensitivity to unmeasured confounding: part 1</title>
      <link>https://www.rdatagen.net/post/what-does-it-mean-if-findings-are-sensitive-to-unmeasured-confounding/</link>
      <pubDate>Wed, 02 Jan 2019 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/what-does-it-mean-if-findings-are-sensitive-to-unmeasured-confounding/</guid>
      <description>&lt;p&gt;Principled causal inference methods can be used to compare the effects of different exposures or treatments we have observed in non-experimental settings. These methods, which include matching (with or without propensity scores), inverse probability weighting, and various g-methods, help us create comparable groups to simulate a randomized experiment. All of these approaches rely on a key assumption of &lt;em&gt;no unmeasured confounding&lt;/em&gt;. The problem is, short of subject matter knowledge, there is no way to test this assumption empirically.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Parallel processing to add a little zip to power simulations (and other replication studies)</title>
      <link>https://www.rdatagen.net/post/parallel-processing-to-add-a-little-zip-to-power-simulations/</link>
      <pubDate>Mon, 10 Dec 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/parallel-processing-to-add-a-little-zip-to-power-simulations/</guid>
      <description>&lt;p&gt;It’s always nice to be able to speed things up a bit. My &lt;a href=&#34;https://www.rdatagen.net/post/first-blog-entry/&#34;&gt;first blog post ever&lt;/a&gt; described an approach using &lt;code&gt;Rcpp&lt;/code&gt; to make huge improvements in a particularly intensive computational process. Here, I want to show how simple it is to speed things up by using the R package &lt;code&gt;parallel&lt;/code&gt; and its function &lt;code&gt;mclapply&lt;/code&gt;. I’ve been using this function more and more, so I want to explicitly demonstrate it in case any one is wondering.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Horses for courses, or to each model its own (causal effect)</title>
      <link>https://www.rdatagen.net/post/different-models-estimate-different-causal-effects-part-ii/</link>
      <pubDate>Wed, 28 Nov 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/different-models-estimate-different-causal-effects-part-ii/</guid>
      <description>&lt;p&gt;In my previous &lt;a href=&#34;https://www.rdatagen.net/post/generating-data-to-explore-the-myriad-causal-effects/&#34;&gt;post&lt;/a&gt;, I described a (relatively) simple way to simulate observational data in order to compare different methods to estimate the causal effect of some exposure or treatment on an outcome. The underlying data generating process (DGP) included a possibly unmeasured confounder and an instrumental variable. (If you haven’t already, you should probably take a quick &lt;a href=&#34;https://www.rdatagen.net/post/generating-data-to-explore-the-myriad-causal-effects/&#34;&gt;look&lt;/a&gt;.)&lt;/p&gt;&#xA;&lt;p&gt;A key point in considering causal effect estimation is that the average causal effect depends on the individuals included in the average. If we are talking about the causal effect for the population - that is, comparing the average outcome if &lt;em&gt;everyone&lt;/em&gt; in the population received treatment against the average outcome if &lt;em&gt;no one&lt;/em&gt; in the population received treatment - then we are interested in the average causal effect (ACE).&lt;/p&gt;</description>
    </item>
    <item>
      <title>Generating data to explore the myriad causal effects that can be estimated in observational data analysis</title>
      <link>https://www.rdatagen.net/post/generating-data-to-explore-the-myriad-causal-effects/</link>
      <pubDate>Tue, 20 Nov 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/generating-data-to-explore-the-myriad-causal-effects/</guid>
      <description>&lt;p&gt;I’ve been inspired by two recent talks describing the challenges of using instrumental variable (IV) methods. IV methods are used to estimate the causal effects of an exposure or intervention when there is unmeasured confounding. This estimated causal effect is very specific: the complier average causal effect (CACE). But, the CACE is just one of several possible causal estimands that we might be interested in. For example, there’s the average causal effect (ACE) that represents a population average (not just based the subset of compliers). Or there’s the average causal effect for the exposed or treated (ACT) that allows for the fact that the exposed could be different from the unexposed.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Causal mediation estimation measures the unobservable</title>
      <link>https://www.rdatagen.net/post/causal-mediation/</link>
      <pubDate>Tue, 06 Nov 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/causal-mediation/</guid>
      <description>&lt;p&gt;I put together a series of demos for a group of epidemiology students who are studying causal mediation analysis. Since mediation analysis is not always so clear or intuitive, I thought, of course, that going through some examples of simulating data for this process could clarify things a bit.&lt;/p&gt;&#xA;&lt;p&gt;Quite often we are interested in understanding the relationship between an exposure or intervention on an outcome. Does exposure &lt;span class=&#34;math inline&#34;&gt;\(A\)&lt;/span&gt; (could be randomized or not) have an effect on outcome &lt;span class=&#34;math inline&#34;&gt;\(Y\)&lt;/span&gt;?&lt;/p&gt;</description>
    </item>
    <item>
      <title>Cross-over study design with a major constraint</title>
      <link>https://www.rdatagen.net/post/when-the-research-question-doesn-t-fit-nicely-into-a-standard-study-design/</link>
      <pubDate>Tue, 23 Oct 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/when-the-research-question-doesn-t-fit-nicely-into-a-standard-study-design/</guid>
      <description>&lt;p&gt;Every new study presents its own challenges. (I would have to say that one of the great things about being a biostatistician is the immense variety of research questions that I get to wrestle with.) Recently, I was approached by a group of researchers who wanted to evaluate an intervention. Actually, they had two, but the second one was a minor tweak added to the first. They were trying to figure out how to design the study to answer two questions: (1) is intervention &lt;span class=&#34;math inline&#34;&gt;\(A\)&lt;/span&gt; better than doing nothing and (2) is &lt;span class=&#34;math inline&#34;&gt;\(A^+\)&lt;/span&gt;, the slightly augmented version of &lt;span class=&#34;math inline&#34;&gt;\(A\)&lt;/span&gt;, better than just &lt;span class=&#34;math inline&#34;&gt;\(A\)&lt;/span&gt;?&lt;/p&gt;</description>
    </item>
    <item>
      <title>In regression, we assume noise is independent of all measured predictors. What happens if it isn&#39;t?</title>
      <link>https://www.rdatagen.net/post/linear-regression-models-assume-noise-is-independent/</link>
      <pubDate>Tue, 09 Oct 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/linear-regression-models-assume-noise-is-independent/</guid>
      <description>&lt;p&gt;A number of key assumptions underlie the linear regression model - among them linearity and normally distributed noise (error) terms with constant variance In this post, I consider an additional assumption: the unobserved noise is uncorrelated with any covariates or predictors in the model.&lt;/p&gt;&#xA;&lt;p&gt;In this simple model:&lt;/p&gt;&#xA;&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[Y_i = \beta_0 + \beta_1X_i + e_i,\]&lt;/span&gt;&lt;/p&gt;&#xA;&lt;p&gt;&lt;span class=&#34;math inline&#34;&gt;\(Y_i\)&lt;/span&gt; has both a structural and stochastic (random) component. The structural component is the linear relationship of &lt;span class=&#34;math inline&#34;&gt;\(Y\)&lt;/span&gt; with &lt;span class=&#34;math inline&#34;&gt;\(X\)&lt;/span&gt;. The random element is often called the &lt;span class=&#34;math inline&#34;&gt;\(error\)&lt;/span&gt; term, but I prefer to think of it as &lt;span class=&#34;math inline&#34;&gt;\(noise\)&lt;/span&gt;. &lt;span class=&#34;math inline&#34;&gt;\(e_i\)&lt;/span&gt; is not measuring something that has gone awry, but rather it is variation emanating from some unknown, unmeasurable source or sources for each individual &lt;span class=&#34;math inline&#34;&gt;\(i\)&lt;/span&gt;. It represents everything we haven’t been able to measure.&lt;/p&gt;</description>
    </item>
    <item>
      <title>simstudy update: improved correlated binary outcomes</title>
      <link>https://www.rdatagen.net/post/simstudy-update-to-version-0-1-10/</link>
      <pubDate>Tue, 25 Sep 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/simstudy-update-to-version-0-1-10/</guid>
      <description>&lt;p&gt;An updated version of the &lt;code&gt;simstudy&lt;/code&gt; package (0.1.10) is now available on &lt;a href=&#34;https://cran.r-project.org/web/packages/simstudy/index.html&#34;&gt;CRAN&lt;/a&gt;. The impetus for this release was a series of requests about generating correlated binary outcomes. In the last &lt;a href=&#34;https://www.rdatagen.net/post/binary-beta-beta-binomial/&#34;&gt;post&lt;/a&gt;, I described a beta-binomial data generating process that uses the recently added beta distribution. In addition to that update, I’ve added functionality to &lt;code&gt;genCorGen&lt;/code&gt; and &lt;code&gt;addCorGen&lt;/code&gt;, functions which generate correlated data from non-Gaussian or normally distributed data such as Poisson, Gamma, and binary data. Most significantly, there is a newly implemented algorithm based on the work of &lt;a href=&#34;https://www.tandfonline.com/doi/abs/10.1080/00031305.1991.10475828&#34;&gt;Emrich &amp;amp; Piedmonte&lt;/a&gt;, which I mentioned the last time around.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Binary, beta, beta-binomial</title>
      <link>https://www.rdatagen.net/post/binary-beta-beta-binomial/</link>
      <pubDate>Tue, 11 Sep 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/binary-beta-beta-binomial/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;I’ve been working on updates for the &lt;a href=&#34;http://www.rdatagen.net/page/simstudy/&#34;&gt;&lt;code&gt;simstudy&lt;/code&gt;&lt;/a&gt; package. In the past few weeks, a couple of folks independently reached out to me about generating correlated binary data. One user was not impressed by the copula algorithm that is already implemented. I’ve added an option to use an algorithm developed by &lt;a href=&#34;https://www.tandfonline.com/doi/abs/10.1080/00031305.1991.10475828&#34;&gt;Emrich and Piedmonte&lt;/a&gt; in 1991, and will be incorporating that option soon in the functions &lt;code&gt;genCorGen&lt;/code&gt; and &lt;code&gt;addCorGen&lt;/code&gt;. I’ll write about that change some point soon.&lt;/p&gt;</description>
    </item>
    <item>
      <title>The power of stepped-wedge designs</title>
      <link>https://www.rdatagen.net/post/alternatives-to-stepped-wedge-designs/</link>
      <pubDate>Tue, 28 Aug 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/alternatives-to-stepped-wedge-designs/</guid>
      <description>&lt;p&gt;Just before heading out on vacation last month, I put up a &lt;a href=&#34;https://www.rdatagen.net/post/by-vs-within/&#34;&gt;post&lt;/a&gt; that purported to compare stepped-wedge study designs with more traditional cluster randomized trials. Either because I rushed or was just lazy, I didn’t exactly do what I set out to do. I &lt;em&gt;did&lt;/em&gt; confirm that a multi-site randomized clinical trial can be more efficient than a cluster randomized trial when there is variability across clusters. (I compared randomizing within a cluster with randomization by cluster.) But, this really had nothing to with stepped-wedge designs.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Multivariate ordinal categorical data generation</title>
      <link>https://www.rdatagen.net/post/multivariate-ordinal-categorical-data-generation/</link>
      <pubDate>Wed, 15 Aug 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/multivariate-ordinal-categorical-data-generation/</guid>
      <description>&lt;p&gt;An economist contacted me about the ability of &lt;code&gt;simstudy&lt;/code&gt; to generate correlated ordinal categorical outcomes. He is trying to generate data as an aide to teaching cost-effectiveness analysis, and is hoping to simulate responses to a quality-of-life survey instrument, the EQ-5D. The particular instrument has five questions related to mobility, self-care, activities, pain, and anxiety. Each item has three possible responses: (1) no problems, (2) some problems, and (3) a lot of problems. Although the instrument has been designed so that each item is orthogonal (independent) from the others, it is impossible to avoid correlation. So, in generating (and analyzing) these kinds of data, it is important to take this into consideration.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Randomize by, or within, cluster?</title>
      <link>https://www.rdatagen.net/post/by-vs-within/</link>
      <pubDate>Thu, 19 Jul 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/by-vs-within/</guid>
      <description>&lt;p&gt;I am involved with a &lt;em&gt;stepped-wedge&lt;/em&gt; designed study that is exploring whether we can improve care for patients with end-stage disease who show up in the emergency room. The plan is to train nurses and physicians in palliative care. (A while ago, I &lt;a href=&#34;https://www.rdatagen.net/post/using-simulation-for-power-analysis-an-example/&#34;&gt;described&lt;/a&gt; what the stepped wedge design is.)&lt;/p&gt;&#xA;&lt;p&gt;Under this design, 33 sites around the country will receive the training at some point, which is no small task (and fortunately as the statistician, this is a part of the study I have little involvement). After hearing about this ambitious plan, a colleague asked why we didn’t just randomize half the sites to the intervention and conduct a more standard cluster randomized trial, where a site would either get the training or not. I quickly simulated some data to see what we would give up (or gain) if we had decided to go that route. (It is actually a moot point, since there would be no way to simultaneously train 16 or so sites, which is why we opted for the stepped-wedge design in the first place.)&lt;/p&gt;</description>
    </item>
    <item>
      <title>How the odds ratio confounds: a brief study in a few colorful figures
</title>
      <link>https://www.rdatagen.net/post/log-odds/</link>
      <pubDate>Tue, 10 Jul 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/log-odds/</guid>
      <description>&lt;p&gt;The odds ratio always confounds: while it may be constant across different groups or clusters, the risk ratios or risk differences across those groups may vary quite substantially. This makes it really hard to interpret an effect. And then there is inconsistency between marginal and conditional odds ratios, a topic I seem to be visiting frequently, most recently last &lt;a href=&#34;https://www.rdatagen.net/post/mixed-effect-models-vs-gee/&#34;&gt;month&lt;/a&gt;.&lt;/p&gt;&#xA;&lt;p&gt;My aim here is to generate a few figures that might highlight some of these issues.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Re-referencing factor levels to estimate standard errors when there is interaction turns out to be a really simple solution</title>
      <link>https://www.rdatagen.net/post/re-referencing-to-estimate-effects-when-there-is-interaction/</link>
      <pubDate>Tue, 26 Jun 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/re-referencing-to-estimate-effects-when-there-is-interaction/</guid>
      <description>&lt;p&gt;Maybe this should be filed under topics that are so obvious that it is not worth writing about. But, I hate to let a good simulation just sit on my computer. I was recently working on a paper investigating the relationship of emotion knowledge (EK) in very young kids with academic performance a year or two later. The idea is that kids who are more emotionally intelligent might be better prepared to learn. My collaborator suspected that the relationship between EK and academics would be different for immigrant and non-immigrant children, so we agreed that this would be a key focus of the analysis.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Late anniversary edition redux: conditional vs marginal models for clustered data
</title>
      <link>https://www.rdatagen.net/post/mixed-effect-models-vs-gee/</link>
      <pubDate>Wed, 13 Jun 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/mixed-effect-models-vs-gee/</guid>
      <description>&lt;p&gt;This afternoon, I was looking over some simulations I plan to use in an upcoming lecture on multilevel models. I created these examples a while ago, before I started this blog. But since it was just about a year ago that I first wrote about this topic (and started the blog), I thought I’d post this now to mark the occasion.&lt;/p&gt;&#xA;&lt;p&gt;The code below provides another way to visualize the difference between marginal and conditional logistic regression models for clustered data (see &lt;a href=&#34;https://www.rdatagen.net/post/marginal-v-conditional/&#34;&gt;here&lt;/a&gt; for an earlier post that discusses in greater detail some of the key issues raised here.) The basic idea is that both models for a binary outcome are valid, but they provide estimates for different quantities.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A little function to help generate ICCs in simple clustered data</title>
      <link>https://www.rdatagen.net/post/a-little-function-to-help-generate-iccs-in-simple-clustered-data/</link>
      <pubDate>Thu, 24 May 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/a-little-function-to-help-generate-iccs-in-simple-clustered-data/</guid>
      <description>&lt;p&gt;In health services research, experiments are often conducted at the provider or site level rather than the patient level. However, we might still be interested in the outcome at the patient level. For example, we could be interested in understanding the effect of a training program for physicians on their patients. It would be very difficult to randomize patients to be exposed or not to the training if a group of patients all see the same doctor. So the experiment is set up so that only some doctors get the training and others serve as the control; we still compare the outcome at the patient level.&lt;/p&gt;</description>
    </item>
    <item>
      <title>How efficient are multifactorial experiments?</title>
      <link>https://www.rdatagen.net/post/so-how-efficient-are-multifactorial-experiments-part/</link>
      <pubDate>Wed, 02 May 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/so-how-efficient-are-multifactorial-experiments-part/</guid>
      <description>&lt;p&gt;I &lt;a href=&#34;https://www.rdatagen.net/post/testing-many-interventions-in-a-single-experiment/&#34;&gt;recently described&lt;/a&gt; why we might want to conduct a multi-factorial experiment, and I alluded to the fact that this approach can be quite efficient. It is efficient in the sense that it is possible to test simultaneously the impact of &lt;em&gt;multiple&lt;/em&gt; interventions using an overall sample size that would be required to test a &lt;em&gt;single&lt;/em&gt; intervention in a more traditional RCT. I demonstrate that here, first with a continuous outcome and then with a binary outcome.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Testing multiple interventions in a single experiment</title>
      <link>https://www.rdatagen.net/post/testing-many-interventions-in-a-single-experiment/</link>
      <pubDate>Thu, 19 Apr 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/testing-many-interventions-in-a-single-experiment/</guid>
      <description>&lt;p&gt;A reader recently inquired about functions in &lt;code&gt;simstudy&lt;/code&gt; that could generate data for a balanced multi-factorial design. I had to report that nothing really exists. A few weeks later, a colleague of mine asked if I could help estimate the appropriate sample size for a study that plans to use a multi-factorial design to choose among a set of interventions to improve rates of smoking cessation. In the course of exploring this, I realized it would be super helpful if the function suggested by the reader actually existed. So, I created &lt;code&gt;genMultiFac&lt;/code&gt;. And since it is now written (though not yet implemented), I thought I’d share some of what I learned (and maybe not yet learned) about this innovative study design.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Exploring the underlying theory of the chi-square test through simulation - part 2</title>
      <link>https://www.rdatagen.net/post/a-little-intuition-and-simulation-behind-the-chi-square-test-of-independence-part-2/</link>
      <pubDate>Sun, 25 Mar 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/a-little-intuition-and-simulation-behind-the-chi-square-test-of-independence-part-2/</guid>
      <description>&lt;p&gt;In the last &lt;a href=&#34;https://www.rdatagen.net/post/a-little-intuition-and-simulation-behind-the-chi-square-test-of-independence/&#34;&gt;post&lt;/a&gt;, I tried to provide a little insight into the chi-square test. In particular, I used simulation to demonstrate the relationship between the Poisson distribution of counts and the chi-squared distribution. The key point in that post was the role conditioning plays in that relationship by reducing variance.&lt;/p&gt;&#xA;&lt;p&gt;To motivate some of the key issues, I talked a bit about recycling. I asked you to imagine a set of bins placed in different locations to collect glass bottles. I will stick with this scenario, but instead of just glass bottle bins, we now also have cardboard, plastic, and metal bins at each location. In this expanded scenario, we are interested in understanding the relationship between location and material. A key question that we might ask: is the distribution of materials the same across the sites? (Assume we are still just counting items and not considering volume or weight.)&lt;/p&gt;</description>
    </item>
    <item>
      <title>Exploring the underlying theory of the chi-square test through simulation - part 1</title>
      <link>https://www.rdatagen.net/post/a-little-intuition-and-simulation-behind-the-chi-square-test-of-independence/</link>
      <pubDate>Sun, 18 Mar 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/a-little-intuition-and-simulation-behind-the-chi-square-test-of-independence/</guid>
      <description>&lt;p&gt;Kids today are so sophisticated (at least they are in New York City, where I live). While I didn’t hear about the chi-square test of independence until my first stint in graduate school, they’re already talking about it in high school. When my kids came home and started talking about it, I did what I usually do when they come home asking about a new statistical concept. I opened up R and started generating some data. Of course, they rolled their eyes, but when the evening was done, I had something that might illuminate some of what underlies the theory of this ubiquitous test.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Another reason to be careful about what you control for</title>
      <link>https://www.rdatagen.net/post/another-reason-to-be-careful-about-what-you-control-for/</link>
      <pubDate>Wed, 07 Mar 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/another-reason-to-be-careful-about-what-you-control-for/</guid>
      <description>&lt;p&gt;Modeling data without any underlying causal theory can sometimes lead you down the wrong path, particularly if you are interested in understanding the &lt;em&gt;way&lt;/em&gt; things work rather than making &lt;em&gt;predictions.&lt;/em&gt; A while back, I &lt;a href=&#34;https://www.rdatagen.net/post/be-careful/&#34;&gt;described&lt;/a&gt; what can go wrong when you control for a mediator when you are interested in an exposure and an outcome. Here, I describe the potential biases that are introduced when you inadvertently control for a variable that turns out to be a &lt;strong&gt;&lt;em&gt;collider&lt;/em&gt;&lt;/strong&gt;.&lt;/p&gt;</description>
    </item>
    <item>
      <title>“I have to randomize by cluster. Is it OK if I only have 6 sites?&#34;</title>
      <link>https://www.rdatagen.net/post/i-have-to-randomize-by-site-is-it-ok-if-i-only-have-6/</link>
      <pubDate>Wed, 21 Feb 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/i-have-to-randomize-by-site-is-it-ok-if-i-only-have-6/</guid>
      <description>&lt;p&gt;The answer is probably no, because there is a not-so-low chance (perhaps considerably higher than 5%) you will draw the wrong conclusions from the study. I have heard variations on this question not so infrequently, so I thought it would be useful (of course) to do a few quick simulations to see what happens when we try to conduct a study under these conditions. (Another question I get every so often, after a study has failed to find an effect: “can we get a post-hoc estimate of the power?” I was all set to post on the issue, but then I found &lt;a href=&#34;http://daniellakens.blogspot.com/2014/12/observed-power-and-what-to-do-if-your.html&#34;&gt;this&lt;/a&gt;, which does a really good job of explaining why this is not a very useful exercise.) But, back to the question at hand.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Have you ever asked yourself, &#34;how should I approach the classic pre-post analysis?&#34;</title>
      <link>https://www.rdatagen.net/post/thinking-about-the-run-of-the-mill-pre-post-analysis/</link>
      <pubDate>Sun, 28 Jan 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/thinking-about-the-run-of-the-mill-pre-post-analysis/</guid>
      <description>&lt;p&gt;Well, maybe you haven’t, but this seems to come up all the time. An investigator wants to assess the effect of an intervention on a outcome. Study participants are randomized either to receive the intervention (could be a new drug, new protocol, behavioral intervention, whatever) or treatment as usual. For each participant, the outcome measure is recorded at baseline - this is the &lt;em&gt;pre&lt;/em&gt; in pre/post analysis. The intervention is delivered (or not, in the case of the control group), some time passes, and the outcome is measured a second time. This is our &lt;em&gt;post&lt;/em&gt;. The question is, how should we analyze this study to draw conclusions about the intervention’s effect on the outcome?&lt;/p&gt;</description>
    </item>
    <item>
      <title>Importance sampling adds an interesting twist to Monte Carlo simulation</title>
      <link>https://www.rdatagen.net/post/importance-sampling-adds-a-little-excitement-to-monte-carlo-simulation/</link>
      <pubDate>Thu, 18 Jan 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/importance-sampling-adds-a-little-excitement-to-monte-carlo-simulation/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;I’m contemplating the idea of teaching a course on simulation next fall, so I have been exploring various topics that I might include. (If anyone has great ideas either because you have taught such a course or taken one, definitely drop me a note.) Monte Carlo (MC) simulation is an obvious one. I like the idea of talking about &lt;em&gt;importance sampling&lt;/em&gt;, because it sheds light on the idea that not all MC simulations are created equally. I thought I’d do a brief post to share some code I put together that demonstrates MC simulation generally, and shows how importance sampling can be an improvement.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Simulating a cost-effectiveness analysis to highlight new functions for generating correlated data</title>
      <link>https://www.rdatagen.net/post/generating-correlated-data-for-a-simulated-cost-effectiveness-analysis/</link>
      <pubDate>Mon, 08 Jan 2018 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/generating-correlated-data-for-a-simulated-cost-effectiveness-analysis/</guid>
      <description>&lt;p&gt;My dissertation work (which I only recently completed - in 2012 - even though I am not exactly young, a whole story on its own) focused on inverse probability weighting methods to estimate a causal cost-effectiveness model. I don’t really do any cost-effectiveness analysis (CEA) anymore, but it came up very recently when some folks in the Netherlands contacted me about using &lt;code&gt;simstudy&lt;/code&gt; to generate correlated (and clustered) data to compare different approaches to estimating cost-effectiveness. As part of this effort, I developed two more functions in simstudy that allow users to generate correlated data drawn from different types of distributions. Earlier I had created the &lt;code&gt;CorGen&lt;/code&gt; functions to generate multivariate data from a single distribution – e.g. multivariate gamma. Now, with the new &lt;code&gt;CorFlex&lt;/code&gt; functions (&lt;code&gt;genCorFlex&lt;/code&gt; and &lt;code&gt;addCorFlex&lt;/code&gt;), users can mix and match distributions. The new version of simstudy is not yet up on CRAN, but is available for download from my &lt;a href=&#34;https://github.com/kgoldfeld/simstudy&#34;&gt;github&lt;/a&gt; site. If you use RStudio, you can install using &lt;code&gt;devtools::install.github(&amp;quot;kgoldfeld/simstudy&amp;quot;)&lt;/code&gt;. [Update: &lt;code&gt;simstudy&lt;/code&gt; version 0.1.8 is now available on &lt;a href=&#34;https://cran.rstudio.com/web/packages/simstudy/&#34;&gt;CRAN&lt;/a&gt;.]&lt;/p&gt;</description>
    </item>
    <item>
      <title>When there&#39;s a fork in the road, take it. Or, taking a look at marginal structural models.</title>
      <link>https://www.rdatagen.net/post/when-a-covariate-is-a-confounder-and-a-mediator/</link>
      <pubDate>Mon, 11 Dec 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/when-a-covariate-is-a-confounder-and-a-mediator/</guid>
      <description>&lt;p&gt;I am going to cut right to the chase, since this is the third of three posts related to confounding and weighting, and it’s kind of a long one. (If you want to catch up, the first two are &lt;a href=&#34;https://www.rdatagen.net/post/potential-outcomes-confounding/&#34;&gt;here&lt;/a&gt; and &lt;a href=&#34;https://www.rdatagen.net/post/inverse-probability-weighting-when-the-outcome-is-binary/&#34;&gt;here&lt;/a&gt;.) My aim with these three posts is to provide a basic explanation of the &lt;em&gt;marginal structural model&lt;/em&gt; (MSM) and how we should interpret the estimates. This is obviously a very rich topic with a vast literature, so if you remain interested in the topic, I recommend checking out this (as of yet unpublished) &lt;a href=&#34;https://www.hsph.harvard.edu/miguel-hernan/causal-inference-book/&#34;&gt;text book&lt;/a&gt; by Hernán &amp;amp; Robins for starters.&lt;/p&gt;</description>
    </item>
    <item>
      <title>When you use inverse probability weighting for estimation, what are the weights actually doing?</title>
      <link>https://www.rdatagen.net/post/inverse-probability-weighting-when-the-outcome-is-binary/</link>
      <pubDate>Mon, 04 Dec 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/inverse-probability-weighting-when-the-outcome-is-binary/</guid>
      <description>&lt;p&gt;Towards the end of &lt;a href=&#34;https://www.rdatagen.net/post/potential-outcomes-confounding/&#34;&gt;Part 1&lt;/a&gt; of this short series on confounding, IPW, and (hopefully) marginal structural models, I talked a little bit about the fact that inverse probability weighting (IPW) can provide unbiased estimates of marginal causal effects in the context of confounding just as more traditional regression models like OLS can. I used an example based on a normally distributed outcome. Now, that example wasn’t super interesting, because in the case of a linear model with homogeneous treatment effects (i.e. no interaction), the marginal causal effect is the same as the conditional effect (that is, conditional on the confounders.) There was no real reason to use IPW in that example - I just wanted to illustrate that the estimates looked reasonable.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Characterizing the variance for clustered data that are Gamma distributed</title>
      <link>https://www.rdatagen.net/post/icc-for-gamma-distribution/</link>
      <pubDate>Mon, 27 Nov 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/icc-for-gamma-distribution/</guid>
      <description>&lt;p&gt;Way back when I was studying algebra and wrestling with one word problem after another (I think now they call them story problems), I complained to my father. He laughed and told me to get used to it. “Life is one big word problem,” is how he put it. Well, maybe one could say any statistical analysis is really just some form of multilevel data analysis, whether we treat it that way or not.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Visualizing how confounding biases estimates of population-wide (or marginal) average causal effects</title>
      <link>https://www.rdatagen.net/post/potential-outcomes-confounding/</link>
      <pubDate>Thu, 16 Nov 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/potential-outcomes-confounding/</guid>
      <description>&lt;p&gt;When we are trying to assess the effect of an exposure or intervention on an outcome, confounding is an ever-present threat to our ability to draw the proper conclusions. My goal (starting here and continuing in upcoming posts) is to think a bit about how to characterize confounding in a way that makes it possible to literally see why improperly estimating intervention effects might lead to bias.&lt;/p&gt;&#xA;&lt;div id=&#34;confounding-potential-outcomes-and-causal-effects&#34; class=&#34;section level3&#34;&gt;&#xA;&lt;h3&gt;Confounding, potential outcomes, and causal effects&lt;/h3&gt;&#xA;&lt;p&gt;Typically, we think of a confounder as a factor that influences &lt;em&gt;both&lt;/em&gt; exposure &lt;em&gt;and&lt;/em&gt; outcome. If we ignore the confounding factor in estimating the effect of an exposure, we can easily over- or underestimate the size of the effect due to the exposure. If sicker patients are more likely than healthier patients to take a particular drug, the relatively poor outcomes of those who took the drug may be due to the initial health status rather than the drug.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A simstudy update provides an excuse to generate and display Likert-type data</title>
      <link>https://www.rdatagen.net/post/generating-and-displaying-likert-type-data/</link>
      <pubDate>Tue, 07 Nov 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/generating-and-displaying-likert-type-data/</guid>
      <description>&lt;p&gt;I just updated &lt;code&gt;simstudy&lt;/code&gt; to version 0.1.7. It is available on CRAN.&lt;/p&gt;&#xA;&lt;p&gt;To mark the occasion, I wanted to highlight a new function, &lt;code&gt;genOrdCat&lt;/code&gt;, which puts into practice some code that I presented a little while back as part of a discussion of &lt;a href=&#34;https://www.rdatagen.net/post/a-hidden-process-part-2-of-2/&#34;&gt;ordinal logistic regression&lt;/a&gt;. The new function was motivated by a reader/researcher who came across my blog while wrestling with a simulation study. After a little back and forth about how to generate ordinal categorical data, I ended up with a function that might be useful. Here’s a little example that uses the &lt;code&gt;likert&lt;/code&gt; package, which makes plotting Likert-type easy and attractive.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Thinking about different ways to analyze sub-groups in an RCT</title>
      <link>https://www.rdatagen.net/post/sub-group-analysis-in-rct/</link>
      <pubDate>Wed, 01 Nov 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/sub-group-analysis-in-rct/</guid>
      <description>&lt;p&gt;Here’s the scenario: we have an intervention that we think will improve outcomes for a particular population. Furthermore, there are two sub-groups (let’s say defined by which of two medical conditions each person in the population has) and we are interested in knowing if the intervention effect is different for each sub-group.&lt;/p&gt;&#xA;&lt;p&gt;And here’s the question: what is the ideal way to set up a study so that we can assess (1) the intervention effects on the group as a whole, but also (2) the sub-group specific intervention effects?&lt;/p&gt;</description>
    </item>
    <item>
      <title>Who knew likelihood functions could be so pretty?</title>
      <link>https://www.rdatagen.net/post/mle-can-be-pretty/</link>
      <pubDate>Mon, 23 Oct 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/mle-can-be-pretty/</guid>
      <description>&lt;p&gt;I just released a new iteration of &lt;code&gt;simstudy&lt;/code&gt; (version 0.1.6), which fixes a bug or two and adds several spline related routines (available on &lt;a href=&#34;https://cran.r-project.org/web/packages/simstudy/index.html&#34;&gt;CRAN&lt;/a&gt;). The &lt;a href=&#34;https://www.rdatagen.net/post/generating-non-linear-data-using-b-splines/&#34;&gt;previous post&lt;/a&gt; focused on using spline curves to generate data, so I won’t repeat myself here. And, apropos of nothing really - I thought I’d take the opportunity to do a simple simulation to briefly explore the likelihood function. It turns out if we generate lots of them, it can be pretty, and maybe provide a little insight.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Can we use B-splines to generate non-linear data?</title>
      <link>https://www.rdatagen.net/post/generating-non-linear-data-using-b-splines/</link>
      <pubDate>Mon, 16 Oct 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/generating-non-linear-data-using-b-splines/</guid>
      <description>&lt;p&gt;I’m exploring the idea of adding a function or set of functions to the &lt;code&gt;simstudy&lt;/code&gt; package that would make it possible to easily generate non-linear data. One way to do this would be using B-splines. Typically, one uses splines to fit a curve to data, but I thought it might be useful to switch things around a bit to use the underlying splines to generate data. This would facilitate exploring models where we know the assumption of linearity is violated. It would also make it easy to explore spline methods, because as with any other simulated data set, we would know the underlying data generating process.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A minor update to simstudy provides an excuse to talk a bit about the negative binomial and Poisson distributions</title>
      <link>https://www.rdatagen.net/post/a-small-update-to-simstudy-neg-bin/</link>
      <pubDate>Thu, 05 Oct 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/a-small-update-to-simstudy-neg-bin/</guid>
      <description>&lt;p&gt;I just updated &lt;code&gt;simstudy&lt;/code&gt; to version 0.1.5 (available on &lt;a href=&#34;https://cran.r-project.org/web/packages/simstudy/index.html&#34;&gt;CRAN&lt;/a&gt;) so that it now includes several new distributions - &lt;em&gt;exponential&lt;/em&gt;, &lt;em&gt;discrete uniform&lt;/em&gt;, and &lt;em&gt;negative binomial&lt;/em&gt;.&lt;/p&gt;&#xA;&lt;p&gt;As part of the release, I thought I’d explore the negative binomial just a bit, particularly as it relates to the Poisson distribution. The Poisson distribution is a discrete (integer) distribution of outcomes of non-negative values that is often used to describe count outcomes. It is characterized by a mean (or rate) and its variance equals its mean.&lt;/p&gt;</description>
    </item>
    <item>
      <title>CACE closed: EM opens up exclusion restriction (among other things)</title>
      <link>https://www.rdatagen.net/post/em-estimation-of-cace/</link>
      <pubDate>Thu, 28 Sep 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/em-estimation-of-cace/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&lt;link href=&#34;https://www.rdatagen.net/rmarkdown-libs/anchor-sections/anchor-sections.css&#34; rel=&#34;stylesheet&#34; /&gt;&#xA;&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/anchor-sections/anchor-sections.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;This is the third, and probably last, of a series of posts touching on the estimation of &lt;a href=&#34;https://www.rdatagen.net/post/cace-explored/&#34;&gt;complier average causal effects&lt;/a&gt; (CACE) and &lt;a href=&#34;https://www.rdatagen.net/post/simstudy-update-provides-an-excuse-to-talk-a-little-bit-about-the-em-algorithm-and-latent-class/&#34;&gt;latent variable modeling techniques&lt;/a&gt; using an expectation-maximization (EM) algorithm. What follows is a simplistic way to implement an EM algorithm in &lt;code&gt;R&lt;/code&gt; to do principal strata estimation of CACE.&lt;/p&gt;&#xA;&lt;div id=&#34;the-em-algorithm&#34; class=&#34;section level3&#34;&gt;&#xA;&lt;h3&gt;The EM algorithm&lt;/h3&gt;&#xA;&lt;p&gt;In this approach, we assume that individuals fall into one of three possible groups - &lt;em&gt;never-takers&lt;/em&gt;, &lt;em&gt;always-takers&lt;/em&gt;, and &lt;em&gt;compliers&lt;/em&gt; - but we cannot see who is who (except in a couple of cases). For each group, we are interested in estimating the unobserved potential outcomes &lt;span class=&#34;math inline&#34;&gt;\(Y_0\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(Y_1\)&lt;/span&gt; using observed outcome measures of &lt;span class=&#34;math inline&#34;&gt;\(Y\)&lt;/span&gt;. The EM algorithm does this in two steps. The &lt;em&gt;E-step&lt;/em&gt; estimates the missing class membership for each individual, and the &lt;em&gt;M-step&lt;/em&gt; provides maximum likelihood estimates of the group-specific potential outcomes and variation.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A simstudy update provides an excuse to talk a little bit about latent class regression and the EM algorithm</title>
      <link>https://www.rdatagen.net/post/simstudy-update-provides-an-excuse-to-talk-a-little-bit-about-the-em-algorithm-and-latent-class/</link>
      <pubDate>Wed, 20 Sep 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/simstudy-update-provides-an-excuse-to-talk-a-little-bit-about-the-em-algorithm-and-latent-class/</guid>
      <description>&lt;p&gt;I was just going to make a quick announcement to let folks know that I’ve updated the &lt;code&gt;simstudy&lt;/code&gt; package to version 0.1.4 (now available on CRAN) to include functions that allow conversion of columns to factors, creation of dummy variables, and most importantly, specification of outcomes that are more flexibly conditional on previously defined variables. But, as I was coming up with an example that might illustrate the added conditional functionality, I found myself playing with package &lt;code&gt;flexmix&lt;/code&gt;, which uses an Expectation-Maximization (EM) algorithm to estimate latent classes and fit regression models. So, in the end, this turned into a bit more than a brief service announcement.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Complier average causal effect? Exploring what we learn from an RCT with participants who don&#39;t do what they are told</title>
      <link>https://www.rdatagen.net/post/cace-explored/</link>
      <pubDate>Tue, 12 Sep 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/cace-explored/</guid>
      <description>&lt;p&gt;Inspired by a free online &lt;a href=&#34;https://courseplus.jhu.edu/core/index.cfm/go/course.home/coid/8155/&#34;&gt;course&lt;/a&gt; titled &lt;em&gt;Complier Average Causal Effects (CACE) Analysis&lt;/em&gt; and taught by Booil Jo and Elizabeth Stuart (through Johns Hopkins University), I’ve decided to explore the topic a little bit. My goal here isn’t to explain CACE analysis in extensive detail (you should definitely go take the course for that), but to describe the problem generally and then (of course) simulate some data. A plot of the simulated data gives a sense of what we are estimating and assuming. And I end by describing two simple methods to estimate the CACE, which we can compare to the truth (since this is a simulation); next time, I will describe a third way.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Further considerations of a hidden process underlying categorical responses</title>
      <link>https://www.rdatagen.net/post/a-hidden-process-part-2-of-2/</link>
      <pubDate>Tue, 05 Sep 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/a-hidden-process-part-2-of-2/</guid>
      <description>&lt;p&gt;In my &lt;a href=&#34;https://www.rdatagen.net/post/ordinal-regression/&#34;&gt;previous post&lt;/a&gt;, I described a continuous data generating process that can be used to generate discrete, categorical outcomes. In that post, I focused largely on binary outcomes and simple logistic regression just because things are always easier to follow when there are fewer moving parts. Here, I am going to focus on a situation where we have &lt;em&gt;multiple&lt;/em&gt; outcomes, but with a slight twist - these groups of interest can be interpreted in an ordered way. This conceptual latent process can provide another perspective on the models that are typically applied to analyze these types of outcomes.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A hidden process behind binary or other categorical outcomes?</title>
      <link>https://www.rdatagen.net/post/ordinal-regression/</link>
      <pubDate>Mon, 28 Aug 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/ordinal-regression/</guid>
      <description>&lt;p&gt;I was thinking a lot about proportional-odds cumulative logit models last fall while designing a study to evaluate an intervention’s effect on meat consumption. After a fairly extensive pilot study, we had determined that participants can have quite a difficult time recalling precise quantities of meat consumption, so we were forced to move to a categorical response. (This was somewhat unfortunate, because we would not have continuous or even count outcomes, and as a result, might not be able to pick up small changes in behavior.) We opted for a question that was based on 30-day meat consumption: none, 1-3 times per month, 1 time per week, etc. - six groups in total. The question was how best to evaluate effectiveness of the intervention?&lt;/p&gt;</description>
    </item>
    <item>
      <title>Be careful not to control for a post-exposure covariate</title>
      <link>https://www.rdatagen.net/post/be-careful/</link>
      <pubDate>Mon, 21 Aug 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/be-careful/</guid>
      <description>&lt;script src=&#34;https://www.rdatagen.net/rmarkdown-libs/header-attrs/header-attrs.js&#34;&gt;&lt;/script&gt;&#xA;&#xA;&#xA;&lt;p&gt;A researcher was presenting an analysis of the impact various types of childhood trauma might have on subsequent substance abuse in adulthood. Obviously, a very interesting and challenging research question. The statistical model included adjustments for several factors that are plausible confounders of the relationship between trauma and substance use, such as childhood poverty. However, the model also include a measurement for poverty in adulthood - believing it was somehow confounding the relationship of trauma and substance use. A confounder is a common cause of an exposure/treatment and an outcome; it is hard to conceive of adult poverty as a cause of childhood events, even though it might be related to adult substance use (or maybe not). At best, controlling for adult poverty has no impact on the conclusions of the research; less good, though, is the possibility that it will lead to the conclusion that the effect of trauma is less than it actually is.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Should we be concerned about incidence - prevalence bias?</title>
      <link>https://www.rdatagen.net/post/simulating-incidence-prevalence-bias/</link>
      <pubDate>Wed, 09 Aug 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/simulating-incidence-prevalence-bias/</guid>
      <description>&lt;p&gt;Recently, we were planning a study to evaluate the effect of an intervention on outcomes for very sick patients who show up in the emergency department. My collaborator had concerns about a phenomenon that she had observed in other studies that might affect the results - patients measured earlier in the study tend to be sicker than those measured later in the study. This might not be a problem, but in the context of a stepped-wedge study design (see &lt;a href=&#34;https://www.rdatagen.net/post/using-simulation-for-power-analysis-an-example/&#34;&gt;this&lt;/a&gt; for a discussion that touches this type of study design), this could definitely generate biased estimates: when the intervention occurs later in the study (as it does in a stepped-wedge design), the “exposed” and “unexposed” populations could differ, and in turn so could the outcomes. We might confuse an artificial effect as an intervention effect.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Using simulation for power analysis: an example based on a stepped wedge study design</title>
      <link>https://www.rdatagen.net/post/using-simulation-for-power-analysis-an-example/</link>
      <pubDate>Mon, 10 Jul 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/using-simulation-for-power-analysis-an-example/</guid>
      <description>&lt;p&gt;Simulation can be super helpful for estimating power or sample size requirements when the study design is complex. This approach has some advantages over an analytic one (i.e. one based on a formula), particularly the flexibility it affords in setting up the specific assumptions in the planned study, such as time trends, patterns of missingness, or effects of different levels of clustering. A downside is certainly the complexity of writing the code as well as the computation time, which &lt;em&gt;can&lt;/em&gt; be a bit painful. My goal here is to show that at least writing the code need not be overwhelming.&lt;/p&gt;</description>
    </item>
    <item>
      <title>simstudy update: two new functions that generate correlated observations from non-normal distributions</title>
      <link>https://www.rdatagen.net/post/simstudy-update-two-functions-for-correlation/</link>
      <pubDate>Wed, 05 Jul 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/simstudy-update-two-functions-for-correlation/</guid>
      <description>&lt;p&gt;In an earlier &lt;a href=&#34;https://www.rdatagen.net/post/correlated-data-copula/&#34;&gt;post&lt;/a&gt;, I described in a fair amount of detail an algorithm to generate correlated binary or Poisson data. I mentioned that I would be updating &lt;code&gt;simstudy&lt;/code&gt; with functions that would make generating these kind of data relatively painless. Well, I have managed to do that, and the updated package (version 0.1.3) is available for download from &lt;a href=&#34;https://cran.r-project.org/web/packages/simstudy/index.html&#34;&gt;CRAN&lt;/a&gt;. There are now two additional functions to facilitate the generation of correlated data from &lt;em&gt;binomial&lt;/em&gt;, &lt;em&gt;poisson&lt;/em&gt;, &lt;em&gt;gamma&lt;/em&gt;, and &lt;em&gt;uniform&lt;/em&gt; distributions: &lt;code&gt;genCorGen&lt;/code&gt; and &lt;code&gt;addCorGen&lt;/code&gt;. Here’s a brief intro to these functions.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Balancing on multiple factors when the sample is too small to stratify </title>
      <link>https://www.rdatagen.net/post/balancing-when-sample-is-too-small-to-stratify/</link>
      <pubDate>Mon, 26 Jun 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/balancing-when-sample-is-too-small-to-stratify/</guid>
      <description>&lt;p&gt;Ideally, a study that uses randomization provides a balance of characteristics that might be associated with the outcome being studied. This way, we can be more confident that any differences in outcomes between the groups are due to the group assignments and not to differences in characteristics. Unfortunately, randomization does not &lt;em&gt;guarantee&lt;/em&gt; balance, especially with smaller sample sizes. If we want to be certain that groups are balanced with respect to a particular characteristic, we need to do something like stratified randomization.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Copulas and correlated data generation: getting beyond the normal distribution</title>
      <link>https://www.rdatagen.net/post/correlated-data-copula/</link>
      <pubDate>Mon, 19 Jun 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/correlated-data-copula/</guid>
      <description>&lt;p&gt;Using the &lt;code&gt;simstudy&lt;/code&gt; package, it’s possible to generate correlated data from a normal distribution using the function &lt;em&gt;genCorData&lt;/em&gt;. I’ve wanted to extend the functionality so that we can generate correlated data from other sorts of distributions; I thought it would be a good idea to begin with binary and Poisson distributed data, since those come up so frequently in my work. &lt;code&gt;simstudy&lt;/code&gt; can already accommodate more general correlated data, but only in the context of a random effects data generation process. This might not be what we want, particularly if we are interested in explicitly generating data to explore marginal models (such as a GEE model) rather than a conditional random effects model (a topic I explored in my &lt;a href=&#34;https://www.rdatagen.net/post/marginal-v-conditional/&#34;&gt;previous&lt;/a&gt; discussion). The extension can quite easily be done using &lt;em&gt;copulas&lt;/em&gt;.&lt;/p&gt;</description>
    </item>
    <item>
      <title>When marginal and conditional logistic model estimates diverge</title>
      <link>https://www.rdatagen.net/post/marginal-v-conditional/</link>
      <pubDate>Fri, 09 Jun 2017 00:00:00 +0000</pubDate><author>keith.goldfeld@nyumc.org (Keith Goldfeld)</author>
      <guid>https://www.rdatagen.net/post/marginal-v-conditional/</guid>
      <description>&lt;STYLE TYPE=&#34;text/css&#34;&gt;&#xA;&lt;!--&#xA;  td{&#xA;    font-family: Arial; &#xA;    font-size: 9pt;&#xA;    height: 2px;&#xA;    padding:0px;&#xA;    cellpadding=&#34;0&#34;;&#xA;    cellspacing=&#34;0&#34;;&#xA;    text-align: center;&#xA;  }&#xA;  th {&#xA;    font-family: Arial; &#xA;    font-size: 9pt;&#xA;    height: 20px;&#xA;    font-weight: bold;&#xA;    text-align: center;&#xA;  }&#xA;  table { &#xA;    border-spacing: 0px;&#xA;    border-collapse: collapse;&#xA;  }&#xA;---&gt;&#xA;&lt;/STYLE&gt;&#xA;&lt;p&gt;Say we have an intervention that is assigned at a group or cluster level but the outcome is measured at an individual level (e.g. students in different schools, eyes on different individuals). And, say this outcome is binary; that is, something happens, or it doesn’t. (This is important, because none of this is true if the outcome is continuous and close to normally distributed.) If we want to measure the &lt;em&gt;effect&lt;/em&gt; of the intervention - perhaps the risk difference, risk ratio, or odds ratio - it can really matter if we are interested in the &lt;em&gt;marginal&lt;/em&gt; effect or the &lt;em&gt;conditional&lt;/em&gt; effect, because they likely won’t be the same.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
