每日总结
Pandas 数据清洗
Pandas 清洗空值:DataFrame.dropna(axis=0, how='any', thresh=None, subset=None, inplace=False)
例子:import pandas as pd
df = pd.read_csv('property-data.csv')
new_df = df.dropna()
print(new_df.to_string())
Pandas 清洗格式错误数据:
例子:
import pandas as pd
# 第三个日期格式错误
data = {
"Date": ['2020/12/01', '2020/12/02' , '20201226'],
"duration": [50, 40, 45]
}
df = pd.DataFrame(data, index = ["day1", "day2", "day3"])
df['Date'] = pd.to_datetime(df['Date'])
print(df.to_string())
Pandas 清洗重复数据:使用 duplicated() 和 drop_duplicates() 方法