Pandas 基本方法

Pandas 基本方法实例

到目前为止，我们了解了三个Pandas DataStructures以及如何创建它们。由于它在实时数据处理中的重要性，因此我们将主要关注DataFrame对象，并讨论其他一些DataStructures。

方法	描述
axes	返回行轴标签的列表
dtype	返回对象的dtype。
empty	如果Series为空，则返回True。
ndim	根据定义返回基础数据的维数。
size	返回基础数据中的元素数。
values	将Series返回为ndarray。
head()	返回前n行。
tail()	返回最后n行。

接下来我们创建一个Series，并看看上所有列表的属性操作。

Example

import pandas as pd
import numpy as np
# 用100随机数创建一个Series
s = pd.Series(np.random.randn(4))
print(s)

运行结果：

0   0.967853
1  -0.148368
2  -1.395906
3  -1.758394
dtype: float64

axes

返回Series标签的列表

import pandas as pd
import numpy as np
# 用100随机数创建一个Series
s = pd.Series(np.random.randn(4))
print ("The axes are:")
print(s.axes)

运行结果：

The axes are:
[RangeIndex(start=0, stop=4, step=1)]

以上结果是0到5（即[0,1,2,3,4]）。

empty

返回布尔值，说明对象是否为空。True表示对象为空

import pandas as pd
import numpy as np
# 用100随机数创建一个Series
s = pd.Series(np.random.randn(4))
print ("Is the Object empty?")
print(s.empty)

运行结果：

Is the Object empty?
False

ndim

Returns the number of dimensions of the object. By definition, a Series is a 1D data structure, so it returns

import pandas as pd
import numpy as np
# 用4个随机数创建一个Series
s = pd.Series(np.random.randn(4))
print s
print ("The dimensions of the object:")
print(s.ndim)

运行结果：

0   0.175898
1   0.166197
2  -0.609712
3  -1.377000
dtype: float64

The dimensions of the object:
1

size

返回Series的大小（长度）.

import pandas as pd
import numpy as np
# 用4个随机数创建一个Series
s = pd.Series(np.random.randn(2))
print s
print ("The size of the object:")
print(s.size)

运行结果：

0   3.078058
1  -1.207803
dtype: float64

The size of the object:
2

values

以数组形式返回Series数据

import pandas as pd
import numpy as np
# 用4个随机数创建一个Series
s = pd.Series(np.random.randn(4))
print s
print ("The actual data series is:")
print(s.values)

运行结果：

0   1.787373
1  -0.605159
2   0.180477
3  -0.140922
dtype: float64

The actual data series is:
[ 1.78737302 -0.60515881 0.18047664 -0.1409218 ]

Head 和 Tail

要查看Series或DataFrame对象的头尾数据，请使用head() 和tail() 方法。

head() 返回前n行（观察索引值）。默认显示的元素数是5，但是您可以传递自定义数字。

import pandas as pd
import numpy as np
# 用4个随机数创建一个Series
s = pd.Series(np.random.randn(4))
print ("The original series is:")
print s
print ("The first two rows of the data series:")
print(s.head(2))

运行结果：

The original series is:
0   0.720876
1  -0.765898
2   0.479221
3  -0.139547
dtype: float64

The first two rows of the data series:
0   0.720876
1  -0.765898
dtype: float64

tail() 返回最后n行（观察索引值）。默认显示的元素数是5，但是您可以传递自定义数字。

import pandas as pd
import numpy as np
# 用4个随机数创建一个Series
s = pd.Series(np.random.randn(4))
print("The original series is:")
print(s)
print("The last two rows of the data series:")
print(s)tail(2)

运行结果：

The original series is:
0 -0.655091
1 -0.881407
2 -0.608592
3 -2.341413
dtype: float64

The last two rows of the data series:
2 -0.608592
3 -2.341413
dtype: float64

DataFrame 基本功能

现在让我们了解什么是DataFrame基本功能。下表列出了有助于DataFrame基本功能的重要属性或方法。

属性/方法	描述
T	行和列互相转换
axes	返回以行轴标签和列轴标签为唯一成员的列表。
dtypes	返回此对象中的dtypes。
empty	如果NDFrame完全为空[没有项目]，则为true；否则为false。如果任何轴的长度为0。
ndim	轴数/数组尺寸。
shape	返回表示DataFrame维度的元组。
size	NDFrame中的元素数。
values	NDFrame的数字表示。
head()	返回前n行。
tail()	返回最后n行。

下面我们下创建一个DataFrame并查看上述属性的所有操作方式。

Example

import pandas as pd
import numpy as np
# 创建Series字典
d = {'Name':pd.Series(['Tom','James','Ricky','Vin','Steve','Smith','Jack']),
'Age':pd.Series([25,26,25,23,30,29,23]),
'Rating':pd.Series([4.23,3.24,3.98,2.56,3.20,4.6,3.8])}
# 创建一个 DataFrame
df = pd.DataFrame(d)
print ("Our data series is:")
print(df)

运行结果：

Our data series is:
    Age   Name    Rating
0   25    Tom     4.23
1   26    James   3.24
2   25    Ricky   3.98
3   23    Vin     2.56
4   30    Steve   3.20
5   29    Smith   4.60
6   23    Jack    3.80

T (Transpose)

返回DataFrame的转置。行和列将互换。

import pandas as pd
import numpy as np

# 创建Series字典
d = {'Name':pd.Series(['Tom','James','Ricky','Vin','Steve','Smith','Jack']),
'Age':pd.Series([25,26,25,23,30,29,23]),
'Rating':pd.Series([4.23,3.24,3.98,2.56,3.20,4.6,3.8])}
# 创建一个 DataFrame
df = pd.DataFrame(d)
print ("The transpose of the data series is:")
print(df.T)

运行结果：

The transpose of the data series is:
         0     1       2      3      4      5       6
Age      25    26      25     23     30     29      23
Name     Tom   James   Ricky  Vin    Steve  Smith   Jack
Rating   4.23  3.24    3.98   2.56   3.2    4.6     3.8

axes

返回行轴标签和列轴标签的列表。

运行结果：

Row axis labels and column axis labels are:
[RangeIndex(start=0, stop=7, step=1), Index([u'Age', u'Name', u'Rating'],
dtype='object')]

dtypes

返回每一列的数据类型。

运行结果：

The data types of each column are:
Age     int64
Name    object
Rating  float64
dtype: object

empty

返回布尔值，说明对象是否为空；True表示对象为空。

import pandas as pd
import numpy as np

# 创建Series字典
d = {'Name':pd.Series(['Tom','James','Ricky','Vin','Steve','Smith','Jack']),
'Age':pd.Series([25,26,25,23,30,29,23]),
'Rating':pd.Series([4.23,3.24,3.98,2.56,3.20,4.6,3.8])}

# 创建一个 DataFrame
df = pd.DataFrame(d)
print ("Is the object empty?")
print(df.empty)

运行结果：

Is the object empty?
False

ndim

返回对象的数量。根据定义，DataFrame是2D对象。

运行结果：

Our object is:
      Age    Name     Rating
0     25     Tom      4.23
1     26     James    3.24
2     25     Ricky    3.98
3     23     Vin      2.56
4     30     Steve    3.20
5     29     Smith    4.60
6     23     Jack     3.80

The dimension of the object is:
2

shape

返回表示DataFrame维度的元组。元组（a，b），其中a表示行数，b表示列数。

import pandas as pd
import numpy as np

# 创建Series字典
d = {'Name':pd.Series(['Tom','James','Ricky','Vin','Steve','Smith','Jack']),
'Age':pd.Series([25,26,25,23,30,29,23]),
'Rating':pd.Series([4.23,3.24,3.98,2.56,3.20,4.6,3.8])}

# 创建一个 DataFrame
df = pd.DataFrame(d)
print ("Our object is:")
print df
print ("The shape of the object is:")
print(df.shape)

运行结果：

Our object is:
   Age   Name    Rating
0  25    Tom     4.23
1  26    James   3.24
2  25    Ricky   3.98
3  23    Vin     2.56
4  30    Steve   3.20
5  29    Smith   4.60
6  23    Jack    3.80

The shape of the object is:
(7, 3)

size

返回DataFrame中的元素数。

import pandas as pd
import numpy as np

# 创建Series字典
d = {'Name':pd.Series(['Tom','James','Ricky','Vin','Steve','Smith','Jack']),
'Age':pd.Series([25,26,25,23,30,29,23]),
'Rating':pd.Series([4.23,3.24,3.98,2.56,3.20,4.6,3.8])}

# 创建一个 DataFrame
df = pd.DataFrame(d)
print ("Our object is:")
print df
print ("The total number of elements in our object is:")
print(df.size)

运行结果：

Our object is:
    Age   Name    Rating
0   25    Tom     4.23
1   26    James   3.24
2   25    Ricky   3.98
3   23    Vin     2.56
4   30    Steve   3.20
5   29    Smith   4.60
6   23    Jack    3.80

The total number of elements in our object is:
21

values

以NDarray的形式返回DataFrame中的实际数据。

import pandas as pd
import numpy as np

# 创建Series字典
d = {'Name':pd.Series(['Tom','James','Ricky','Vin','Steve','Smith','Jack']),
'Age':pd.Series([25,26,25,23,30,29,23]),
'Rating':pd.Series([4.23,3.24,3.98,2.56,3.20,4.6,3.8])}

# 创建一个 DataFrame
df = pd.DataFrame(d)
print ("Our object is:")
print df
print ("The actual data in our data frame is:")
print(df.values)

运行结果：

Our object is:
    Age   Name    Rating
0   25    Tom     4.23
1   26    James   3.24
2   25    Ricky   3.98
3   23    Vin     2.56
4   30    Steve   3.20
5   29    Smith   4.60
6   23    Jack    3.80
The actual data in our data frame is:
[[25 'Tom' 4.23]
[26 'James' 3.24]
[25 'Ricky' 3.98]
[23 'Vin' 2.56]
[30 'Steve' 3.2]
[29 'Smith' 4.6]
[23 'Jack' 3.8]]

Head & Tail

要查看DataFrame对象的头尾数据，请使用head()和tail()方法。head() 返回前n行（观察索引值）。默认显示的元素数是5，但是您可以传递自定义数字。

import pandas as pd
import numpy as np

# 创建Series字典
d = {'Name':pd.Series(['Tom','James','Ricky','Vin','Steve','Smith','Jack']),
'Age':pd.Series([25,26,25,23,30,29,23]),
'Rating':pd.Series([4.23,3.24,3.98,2.56,3.20,4.6,3.8])}
# 创建一个 DataFrame
df = pd.DataFrame(d)
print ("Our data frame is:")
print df
print ("The first two rows of the data frame is:")
print(df.head(2))

运行结果：

Our data frame is:
    Age   Name    Rating
0   25    Tom     4.23
1   26    James   3.24
2   25    Ricky   3.98
3   23    Vin     2.56
4   30    Steve   3.20
5   29    Smith   4.60
6   23    Jack    3.80

The first two rows of the data frame is:
   Age   Name   Rating
0  25    Tom    4.23
1  26    James  3.24

tail() 返回最后n行（观察索引值）。默认显示的元素数是5，但是您可以传递自定义数字。

import pandas as pd
import numpy as np
# 创建Series字典
d = {'Name':pd.Series(['Tom','James','Ricky','Vin','Steve','Smith','Jack']),
'Age':pd.Series([25,26,25,23,30,29,23]),
'Rating':pd.Series([4.23,3.24,3.98,2.56,3.20,4.6,3.8])}

# 创建一个 DataFrame
df = pd.DataFrame(d)
print ("Our data frame is:")
print df
print ("The last two rows of the data frame is:")
print(df.tail(2))

运行结果：

Our data frame is:
    Age   Name    Rating
0   25    Tom     4.23
1   26    James   3.24
2   25    Ricky   3.98
3   23    Vin     2.56
4   30    Steve   3.20
5   29    Smith   4.60
6   23    Jack    3.80

The last two rows of the data frame is:
    Age   Name    Rating
5   29    Smith    4.6
6   23    Jack     3.8

Python 简介 >>

找工作要求35岁以下，35岁以上的程序员都干什么去了？

长久以来，一直有一个问题困扰着技术人——如何打破“程序员的35岁职业魔咒”，这一天迟早会到来，或早或晚。

或许是选错了行业，程序员薪水虽高，但光鲜的外表下，背后的苦衷只有自己知道。三十多岁本该是一个人事业的黄金期，但技术变化日新月异，行业竞争异常残酷，对一个企业来说，永远有比你更年轻、劳动成本更低的人可以选择，这让你的中年危机提前到来。破局的智慧可以看看这本书！>>

<< Pandas Panel Pandas 描述性统计 >>

昵称：邮箱：