How to subset rows of an R data frame based on duplicate values in a particular column?

R Programming Server Side Programming Programming

Duplication is also a problem that we face during data analysis. We can find the rows with duplicated values in a particular column of an R data frame by using duplicated function inside the subset function. This will return only the duplicate rows based on the column we choose that means the first unique value will not be in the output.

Example

Live Demo

Consider the below data frame:
x1<-1:20
x2<-rpois(20,4)
df1<-data.frame(x1,x2)
df1

Output

Create rows of df1 based on duplicates in column x2 −

Example

subset(df1,duplicated(x2))

Output

Example

Live Demo

y1<-LETTERS[1:20]
y2<-sample(0:5,20,replace=TRUE)
df2<-data.frame(y1,y2)
df2

Output

Create rows of df2 based on duplicates in column y2 −

Example

subset(df2,duplicated(y2))

Output

Nizamuddin Siddiqui

Updated on: 2020-12-05T13:06:08+05:30

11K+ Views

Kickstart Your Career

Get certified by completing the course

Get Started