vault backup: 2026-05-30 16:02:33

This commit is contained in:
Rainyy21
2026-05-30 16:02:33 -04:00
commit 96449f8968
43 changed files with 2837 additions and 0 deletions
+241
View File
@@ -0,0 +1,241 @@
---
title: Introduction to Pandas
tags:
- python
- pandas
- data-science
- data-analysis
category: Lesson
difficulty: Beginner to Intermediate
created: 2026-05-26
---
# Introduction to Pandas
> [!abstract] What is Pandas?
> **Pandas** is a fast, powerful, flexible, and easy-to-use open-source data analysis and manipulation library built on top of the Python programming language.
>
> The name is derived from **"Panel Data"**, an econometrics term for multidimensional structured data sets. It is the foundational library for Data Science, Machine Learning, and Data Analysis in Python, acting as the bridge between raw data files (like CSVs, Excel files, or SQL databases) and numerical/modeling libraries (like NumPy, Scikit-Learn, and PyTorch).
---
## 🏗️ Core Data Structures
Pandas simplifies data manipulation by providing two primary, highly optimized data structures:
```mermaid
graph TD
A[Pandas Data Structures] --> B[Series - 1D]
A --> C[DataFrame - 2D]
B -->|Multiple columns merged| C
C -->|Single column extracted| B
```
### 1. Series (1D)
A **Series** is a one-dimensional array-like object containing an array of data and an associated array of data labels, called its **index**. Think of it as a single column in a spreadsheet.
```python
import pandas as pd
# Creating a Series
temperatures = pd.Series([22.5, 24.0, 19.5, 21.8], name="Temp")
print(temperatures)
```
### 2. DataFrame (2D)
A **DataFrame** represents a tabular, spreadsheet-like data structure containing an ordered collection of columns, each of which can be a different value type (numeric, string, boolean, etc.). It has both a row index and a column index.
| Index | Name | Age | Department |
| :--- | :--- | :--- | :--- |
| **0** | Alice | 28 | Engineering |
| **1** | Bob | 34 | Marketing |
| **2** | Charlie | 22 | HR |
---
## 🛠️ Getting Started & Creating Data
First, make sure you have pandas imported. The universal convention is to alias it as `pd`.
```python
import pandas as pd
import numpy as np
```
### Creating DataFrames Manually
You can easily create DataFrames from dictionaries or lists:
```python
data = {
'Name': ['Alice', 'Bob', 'Charlie', 'David'],
'Age': [25, 30, 35, 40],
'Salary': [70000, 80000, 120000, 90000],
'Department': ['HR', 'Engineering', 'Engineering', 'Finance']
}
df = pd.DataFrame(data)
```
---
## 🔍 Essential Operations (The "Big Five")
### 1. Reading & Writing Data
Pandas supports a wide range of file formats out of the box.
```python
# Reading data
df_csv = pd.read_csv('employees.csv')
df_excel = pd.read_excel('sales.xlsx', sheet_name='Sheet1')
df_sql = pd.read_sql('SELECT * FROM users', database_connection)
# Writing data
df.to_csv('output.csv', index=False) # index=False avoids writing the row numbers
df.to_json('output.json')
```
### 2. Inspecting Your Data
Before doing any analysis, you must understand the shape and types of your data.
```python
df.head(2) # Returns the first 2 rows
df.tail(2) # Returns the last 2 rows
df.info() # Summary of columns, non-null counts, and data types
df.describe() # Generates descriptive statistics for numerical columns
df.shape # Returns (rows, columns) as a tuple
```
### 3. Selection & Filtering
Selecting data is one of the most common tasks. Pandas provides multiple intuitive ways to do this.
#### Column Selection
```python
# Select a single column (returns a Series)
ages = df['Age']
# Select multiple columns (returns a DataFrame)
subset = df[['Name', 'Salary']]
```
#### Row Selection using `.loc` and `.iloc`
* `.loc` is **label-based**: references rows/columns by their row labels or column names.
* `.iloc` is **integer-position-based**: references rows/columns by their 0-indexed positions.
```python
# Select the first row by position
first_row = df.iloc[0]
# Select a cell by row label and column name
val = df.loc[2, 'Name'] # Charlie
```
#### Boolean Indexing (Filtering)
To filter rows based on conditions:
```python
# Filter employees earning more than 85,000
high_earners = df[df['Salary'] > 85000]
# Combine multiple conditions using & (AND) or | (OR)
# Always wrap conditions in parentheses!
eng_seniors = df[(df['Department'] == 'Engineering') & (df['Age'] > 30)]
```
### 4. Data Cleaning
Raw data is rarely perfect. Pandas excels at handling missing values and data type conversions.
```python
# Check for missing values
df.isna().sum()
# Drop rows with missing values
df_clean = df.dropna()
# Fill missing values with a default/placeholder
df['Salary'] = df['Salary'].fillna(df['Salary'].mean())
# Rename columns
df = df.rename(columns={'Name': 'Full Name', 'Age': 'Years'})
```
### 5. Grouping & Aggregation
To summarize data by categories, use the Split-Apply-Combine workflow via `groupby`.
```python
# Calculate the average salary by department
dept_salaries = df.groupby('Department')['Salary'].mean()
# Perform multiple aggregations at once
summary = df.groupby('Department').agg({
'Salary': ['mean', 'min', 'max'],
'Age': 'mean'
})
```
---
## 🎓 Hands-on Practice Walkthrough
Let's walk through a mini-scenario. Imagine we have the following DataFrame of store transactions:
```python
transactions = pd.DataFrame({
'TransactionID': [101, 102, 103, 104, 105],
'Store': ['North', 'South', 'North', 'West', 'South'],
'Amount': [250.50, 150.00, np.nan, 300.25, 450.00],
'ItemCount': [3, 2, 1, 5, 4]
})
```
Let's answer three questions:
### Question 1: Fill missing amounts with the median transaction amount.
```python
median_amount = transactions['Amount'].median() # 275.375
transactions['Amount'] = transactions['Amount'].fillna(median_amount)
```
### Question 2: Find transactions with a total amount greater than 200.
```python
large_tx = transactions[transactions['Amount'] > 200]
```
### Question 3: Find the total revenue (sum of amounts) generated by each store.
```python
store_revenue = transactions.groupby('Store')['Amount'].sum().reset_index()
print(store_revenue)
# Store Amount
# 0 North 525.875
# 1 South 600.00
# 2 West 300.25
```
---
## ⚡ Performance Tips & Best Practices
> [!warning] Don't Loop Over Rows!
> Never write `for index, row in df.iterrows():` unless absolutely necessary. Iteration is extremely slow because it disables pandas' vectorized backend.
>
> **Instead, use Vectorization:**
> ```python
> # ❌ Slow & non-pythonic
> for i in range(len(df)):
> df.loc[i, 'Tax'] = df.loc[i, 'Salary'] * 0.1
>
> # ✅ Fast & vectorised
> df['Tax'] = df['Salary'] * 0.1
> ```
> [!tip] Avoid the SettingWithCopyWarning
> When you slice a DataFrame and then modify it, pandas warns you that you might be editing a temporary copy rather than the original source.
> To prevent this, use `.copy()` when creating a subset you intend to modify:
> ```python
> # ❌ Might trigger SettingWithCopyWarning
> engineering = df[df['Department'] == 'Engineering']
> engineering['Bonus'] = 1000
>
> # ✅ Clean and safe
> engineering = df[df['Department'] == 'Engineering'].copy()
> engineering['Bonus'] = 1000
> ```