pandas列相关性具有统计学意义

马朝斑

2023-03-14

问题内容：

给定一个熊猫数据框df，以获得其列df.1与之间的相关性的最佳方法是什么df.2？

我不希望输出用来计数行NaN，而pandas内置相关性可以。但是我也希望它输出一个pvalue或标准错误，而内置错误则不会。

SciPy 似乎被NaN赶上了，尽管我相信它确实具有重要意义。

数据示例：

     1           2
0    2          NaN
1    NaN         1
2    1           2
3    -4          3
4    1.3         1
5    NaN         NaN

问题答案：

@Shashank提供的答案很好。但是，如果您想使用pure的解决方案pandas，则可能会这样：

import pandas as pd
from pandas.io.data import DataReader
from datetime import datetime
import scipy.stats  as stats


gdp = pd.DataFrame(DataReader("GDP", "fred", start=datetime(1990, 1, 1)))
vix = pd.DataFrame(DataReader("VIXCLS", "fred", start=datetime(1990, 1, 1)))

#Do it with a pandas regression to get the p value from the F-test
df = gdp.merge(vix,left_index=True, right_index=True, how='left')
vix_on_gdp = pd.ols(y=df['VIXCLS'], x=df['GDP'], intercept=True)
print(df['VIXCLS'].corr(df['GDP']), vix_on_gdp.f_stat['p-value'])

结果：

-0.0422917932738 0.851762475093

与统计功能相同的结果：

#Do it with stats functions. 
df_clean = df.dropna()
stats.pearsonr(df_clean['VIXCLS'], df_clean['GDP'])

结果：

  (-0.042291793273791969, 0.85176247509284908)

为了扩展更多的可变项，我给你一个基于丑陋循环的方法：

#Add a third field
oil = pd.DataFrame(DataReader("DCOILWTICO", "fred", start=datetime(1990, 1, 1))) 
df = df.merge(oil,left_index=True, right_index=True, how='left')

#construct two arrays, one of the correlation and the other of the p-vals
rho = df.corr()
pval = np.zeros([df.shape[1],df.shape[1]])
for i in range(df.shape[1]): # rows are the number of rows in the matrix.
    for j in range(df.shape[1]):
        JonI        = pd.ols(y=df.icol(i), x=df.icol(j), intercept=True)
        pval[i,j]  = JonI.f_stat['p-value']

结果：

             GDP    VIXCLS  DCOILWTICO
 GDP         1.000000 -0.042292    0.870251
 VIXCLS     -0.042292  1.000000   -0.004612
 DCOILWTICO  0.870251 -0.004612    1.000000

pval的结果：

 [[  0.00000000e+00   8.51762475e-01   1.11022302e-16]
  [  8.51762475e-01   0.00000000e+00   9.83747425e-01]
  [  1.11022302e-16   9.83747425e-01   0.00000000e+00]]

pandas列相关性具有统计学意义

相关阅读

相关文章

相关问答

相关工具

相关文档