R - Interrater Reliability
04:06 09 Dec 2025

I am working on the psychometric evaluation of a questionnaire. In this questionnaire employees are rating the unit they work in. So I am interested in interrater reliability, that is how much do employees of the same site agree in their ratings. Different sites are rated by different employees and raters are considered random. Based on this paper, I have selected one-way random effects model for absolute agreement (ICC(1,1)).

Here is an example dataframe to show my data structure:

site <- c(rep("A", 7),
          rep("B", 6),
          rep("C", 1),
          rep("D", 1))
rating <- c(6, 5, 4, 6, 4, 4, 5, 
            2, 4, 3, 1, 4, 3,
            2, 6)
df <- data.frame(site, rating)

I looked into psych::ICC for this, but did not figure out how I could apply this function to my data structure.

So my questions are:

  • Main question: What ways are there to calculate the correct ICC?

  • If anyone has a different opinion on which ICC I should use or sees any other trouble in my design, I am happy to learn! For example, most sites were rated by ca. 7 raters, but I also have some sites with just one or two raters (26 sites overall). Is that a problem for ICC validity?

  • Bonus question on missing values: My real data set is much bigger (>200 cases and >200 variables) but also includes some missings where single items were skipped (I can assume they are MAR). Is there a way to calculate an ICC while avoiding listwise deletion and instead using either multiple imputation or full-information maximum likelihood?

Thank you very much!

r reliability psych