keep | Optional , 'first' default, all duplicates are marked True except first one 'last', all duplicates are marked True except last one 'False',all duplicates are marked True
|
import pandas as pd
my_data=pd.Series(['One','Two','Three','Two','Four'])
my_data.duplicated()
Output ( default value of keep is first , keep='first')
0 False
1 False
2 False
3 True
4 False
dtype: bool
with keep='last', duplicate values are marked True except last one
import pandas as pd
my_data=pd.Series(['One','Two','Three','Two','Four'])
my_data.duplicated(keep='last')
0 False
1 True
2 False
3 False
4 False
dtype: bool
with keep=False , all duplicate values are marked True
import pandas as pd
my_data=pd.Series(['One','Two','Three','Two','Four'])
my_data.duplicated(keep=False)
Output
0 False
1 True
2 False
3 True
4 False
dtype: bool
print(my_data[my_data.duplicated()])
Output
3 Two
dtype: object
print(my_data[~my_data.duplicated()])
Output
0 One
1 Two
2 Three
4 Four
dtype: object
We can use unique()
print(my_data.unique())
Output
['One' 'Two' 'Three' 'Four']
Data CleaningAuthor & Instructor at plus2net
I write and maintain practical tutorials on Python, PHP, SQL, JavaScript, HTML, jQuery, and web development at plus2net. The tutorials focus on clear explanations, working examples, and code that readers can test and adapt while learning.