mirror of
https://github.com/Rainyy21/framework_note.git
synced 2026-10-10 23:40:30 -04:00
242 lines
6.9 KiB
Markdown
242 lines
6.9 KiB
Markdown
---
|
|
title: Introduction to Pandas
|
|
tags:
|
|
- python
|
|
- pandas
|
|
- data-science
|
|
- data-analysis
|
|
category: Lesson
|
|
difficulty: Beginner to Intermediate
|
|
created: 2026-05-26
|
|
---
|
|
|
|
# Introduction to Pandas
|
|
|
|
> [!abstract] What is Pandas?
|
|
> **Pandas** is a fast, powerful, flexible, and easy-to-use open-source data analysis and manipulation library built on top of the Python programming language.
|
|
>
|
|
> The name is derived from **"Panel Data"**, an econometrics term for multidimensional structured data sets. It is the foundational library for Data Science, Machine Learning, and Data Analysis in Python, acting as the bridge between raw data files (like CSVs, Excel files, or SQL databases) and numerical/modeling libraries (like NumPy, Scikit-Learn, and PyTorch).
|
|
|
|
---
|
|
|
|
## 🏗️ Core Data Structures
|
|
|
|
Pandas simplifies data manipulation by providing two primary, highly optimized data structures:
|
|
|
|
```mermaid
|
|
graph TD
|
|
A[Pandas Data Structures] --> B[Series - 1D]
|
|
A --> C[DataFrame - 2D]
|
|
B -->|Multiple columns merged| C
|
|
C -->|Single column extracted| B
|
|
```
|
|
|
|
### 1. Series (1D)
|
|
A **Series** is a one-dimensional array-like object containing an array of data and an associated array of data labels, called its **index**. Think of it as a single column in a spreadsheet.
|
|
|
|
```python
|
|
import pandas as pd
|
|
|
|
# Creating a Series
|
|
temperatures = pd.Series([22.5, 24.0, 19.5, 21.8], name="Temp")
|
|
print(temperatures)
|
|
```
|
|
|
|
### 2. DataFrame (2D)
|
|
A **DataFrame** represents a tabular, spreadsheet-like data structure containing an ordered collection of columns, each of which can be a different value type (numeric, string, boolean, etc.). It has both a row index and a column index.
|
|
|
|
| Index | Name | Age | Department |
|
|
| :--- | :--- | :--- | :--- |
|
|
| **0** | Alice | 28 | Engineering |
|
|
| **1** | Bob | 34 | Marketing |
|
|
| **2** | Charlie | 22 | HR |
|
|
|
|
---
|
|
|
|
## 🛠️ Getting Started & Creating Data
|
|
|
|
First, make sure you have pandas imported. The universal convention is to alias it as `pd`.
|
|
|
|
```python
|
|
import pandas as pd
|
|
import numpy as np
|
|
```
|
|
|
|
### Creating DataFrames Manually
|
|
You can easily create DataFrames from dictionaries or lists:
|
|
|
|
```python
|
|
data = {
|
|
'Name': ['Alice', 'Bob', 'Charlie', 'David'],
|
|
'Age': [25, 30, 35, 40],
|
|
'Salary': [70000, 80000, 120000, 90000],
|
|
'Department': ['HR', 'Engineering', 'Engineering', 'Finance']
|
|
}
|
|
|
|
df = pd.DataFrame(data)
|
|
```
|
|
|
|
---
|
|
|
|
## 🔍 Essential Operations (The "Big Five")
|
|
|
|
### 1. Reading & Writing Data
|
|
Pandas supports a wide range of file formats out of the box.
|
|
|
|
```python
|
|
# Reading data
|
|
df_csv = pd.read_csv('employees.csv')
|
|
df_excel = pd.read_excel('sales.xlsx', sheet_name='Sheet1')
|
|
df_sql = pd.read_sql('SELECT * FROM users', database_connection)
|
|
|
|
# Writing data
|
|
df.to_csv('output.csv', index=False) # index=False avoids writing the row numbers
|
|
df.to_json('output.json')
|
|
```
|
|
|
|
### 2. Inspecting Your Data
|
|
Before doing any analysis, you must understand the shape and types of your data.
|
|
|
|
```python
|
|
df.head(2) # Returns the first 2 rows
|
|
df.tail(2) # Returns the last 2 rows
|
|
df.info() # Summary of columns, non-null counts, and data types
|
|
df.describe() # Generates descriptive statistics for numerical columns
|
|
df.shape # Returns (rows, columns) as a tuple
|
|
```
|
|
|
|
### 3. Selection & Filtering
|
|
Selecting data is one of the most common tasks. Pandas provides multiple intuitive ways to do this.
|
|
|
|
#### Column Selection
|
|
```python
|
|
# Select a single column (returns a Series)
|
|
ages = df['Age']
|
|
|
|
# Select multiple columns (returns a DataFrame)
|
|
subset = df[['Name', 'Salary']]
|
|
```
|
|
|
|
#### Row Selection using `.loc` and `.iloc`
|
|
* `.loc` is **label-based**: references rows/columns by their row labels or column names.
|
|
* `.iloc` is **integer-position-based**: references rows/columns by their 0-indexed positions.
|
|
|
|
```python
|
|
# Select the first row by position
|
|
first_row = df.iloc[0]
|
|
|
|
# Select a cell by row label and column name
|
|
val = df.loc[2, 'Name'] # Charlie
|
|
```
|
|
|
|
#### Boolean Indexing (Filtering)
|
|
To filter rows based on conditions:
|
|
|
|
```python
|
|
# Filter employees earning more than 85,000
|
|
high_earners = df[df['Salary'] > 85000]
|
|
|
|
# Combine multiple conditions using & (AND) or | (OR)
|
|
# Always wrap conditions in parentheses!
|
|
eng_seniors = df[(df['Department'] == 'Engineering') & (df['Age'] > 30)]
|
|
```
|
|
|
|
### 4. Data Cleaning
|
|
Raw data is rarely perfect. Pandas excels at handling missing values and data type conversions.
|
|
|
|
```python
|
|
# Check for missing values
|
|
df.isna().sum()
|
|
|
|
# Drop rows with missing values
|
|
df_clean = df.dropna()
|
|
|
|
# Fill missing values with a default/placeholder
|
|
df['Salary'] = df['Salary'].fillna(df['Salary'].mean())
|
|
|
|
# Rename columns
|
|
df = df.rename(columns={'Name': 'Full Name', 'Age': 'Years'})
|
|
```
|
|
|
|
### 5. Grouping & Aggregation
|
|
To summarize data by categories, use the Split-Apply-Combine workflow via `groupby`.
|
|
|
|
```python
|
|
# Calculate the average salary by department
|
|
dept_salaries = df.groupby('Department')['Salary'].mean()
|
|
|
|
# Perform multiple aggregations at once
|
|
summary = df.groupby('Department').agg({
|
|
'Salary': ['mean', 'min', 'max'],
|
|
'Age': 'mean'
|
|
})
|
|
```
|
|
|
|
---
|
|
|
|
## 🎓 Hands-on Practice Walkthrough
|
|
|
|
Let's walk through a mini-scenario. Imagine we have the following DataFrame of store transactions:
|
|
|
|
```python
|
|
transactions = pd.DataFrame({
|
|
'TransactionID': [101, 102, 103, 104, 105],
|
|
'Store': ['North', 'South', 'North', 'West', 'South'],
|
|
'Amount': [250.50, 150.00, np.nan, 300.25, 450.00],
|
|
'ItemCount': [3, 2, 1, 5, 4]
|
|
})
|
|
```
|
|
|
|
Let's answer three questions:
|
|
|
|
### Question 1: Fill missing amounts with the median transaction amount.
|
|
```python
|
|
median_amount = transactions['Amount'].median() # 275.375
|
|
transactions['Amount'] = transactions['Amount'].fillna(median_amount)
|
|
```
|
|
|
|
### Question 2: Find transactions with a total amount greater than 200.
|
|
```python
|
|
large_tx = transactions[transactions['Amount'] > 200]
|
|
```
|
|
|
|
### Question 3: Find the total revenue (sum of amounts) generated by each store.
|
|
```python
|
|
store_revenue = transactions.groupby('Store')['Amount'].sum().reset_index()
|
|
print(store_revenue)
|
|
# Store Amount
|
|
# 0 North 525.875
|
|
# 1 South 600.00
|
|
# 2 West 300.25
|
|
```
|
|
|
|
---
|
|
|
|
## ⚡ Performance Tips & Best Practices
|
|
|
|
> [!warning] Don't Loop Over Rows!
|
|
> Never write `for index, row in df.iterrows():` unless absolutely necessary. Iteration is extremely slow because it disables pandas' vectorized backend.
|
|
>
|
|
> **Instead, use Vectorization:**
|
|
> ```python
|
|
> # ❌ Slow & non-pythonic
|
|
> for i in range(len(df)):
|
|
> df.loc[i, 'Tax'] = df.loc[i, 'Salary'] * 0.1
|
|
>
|
|
> # ✅ Fast & vectorised
|
|
> df['Tax'] = df['Salary'] * 0.1
|
|
> ```
|
|
|
|
> [!tip] Avoid the SettingWithCopyWarning
|
|
> When you slice a DataFrame and then modify it, pandas warns you that you might be editing a temporary copy rather than the original source.
|
|
> To prevent this, use `.copy()` when creating a subset you intend to modify:
|
|
> ```python
|
|
> # ❌ Might trigger SettingWithCopyWarning
|
|
> engineering = df[df['Department'] == 'Engineering']
|
|
> engineering['Bonus'] = 1000
|
|
>
|
|
> # ✅ Clean and safe
|
|
> engineering = df[df['Department'] == 'Engineering'].copy()
|
|
> engineering['Bonus'] = 1000
|
|
> ```
|