Files
framework_note/note/LLM/Introduction to Pandas.md
T
2026-05-30 16:02:33 -04:00

6.9 KiB

title, tags, category, difficulty, created
title tags category difficulty created
Introduction to Pandas
python
pandas
data-science
data-analysis
Lesson Beginner to Intermediate 2026-05-26

Introduction to Pandas

[!abstract] What is Pandas? Pandas is a fast, powerful, flexible, and easy-to-use open-source data analysis and manipulation library built on top of the Python programming language.

The name is derived from "Panel Data", an econometrics term for multidimensional structured data sets. It is the foundational library for Data Science, Machine Learning, and Data Analysis in Python, acting as the bridge between raw data files (like CSVs, Excel files, or SQL databases) and numerical/modeling libraries (like NumPy, Scikit-Learn, and PyTorch).


🏗️ Core Data Structures

Pandas simplifies data manipulation by providing two primary, highly optimized data structures:

graph TD
    A[Pandas Data Structures] --> B[Series - 1D]
    A --> C[DataFrame - 2D]
    B -->|Multiple columns merged| C
    C -->|Single column extracted| B

1. Series (1D)

A Series is a one-dimensional array-like object containing an array of data and an associated array of data labels, called its index. Think of it as a single column in a spreadsheet.

import pandas as pd

# Creating a Series
temperatures = pd.Series([22.5, 24.0, 19.5, 21.8], name="Temp")
print(temperatures)

2. DataFrame (2D)

A DataFrame represents a tabular, spreadsheet-like data structure containing an ordered collection of columns, each of which can be a different value type (numeric, string, boolean, etc.). It has both a row index and a column index.

Index Name Age Department
0 Alice 28 Engineering
1 Bob 34 Marketing
2 Charlie 22 HR

🛠️ Getting Started & Creating Data

First, make sure you have pandas imported. The universal convention is to alias it as pd.

import pandas as pd
import numpy as np

Creating DataFrames Manually

You can easily create DataFrames from dictionaries or lists:

data = {
    'Name': ['Alice', 'Bob', 'Charlie', 'David'],
    'Age': [25, 30, 35, 40],
    'Salary': [70000, 80000, 120000, 90000],
    'Department': ['HR', 'Engineering', 'Engineering', 'Finance']
}

df = pd.DataFrame(data)

🔍 Essential Operations (The "Big Five")

1. Reading & Writing Data

Pandas supports a wide range of file formats out of the box.

# Reading data
df_csv = pd.read_csv('employees.csv')
df_excel = pd.read_excel('sales.xlsx', sheet_name='Sheet1')
df_sql = pd.read_sql('SELECT * FROM users', database_connection)

# Writing data
df.to_csv('output.csv', index=False) # index=False avoids writing the row numbers
df.to_json('output.json')

2. Inspecting Your Data

Before doing any analysis, you must understand the shape and types of your data.

df.head(2)       # Returns the first 2 rows
df.tail(2)       # Returns the last 2 rows
df.info()        # Summary of columns, non-null counts, and data types
df.describe()    # Generates descriptive statistics for numerical columns
df.shape         # Returns (rows, columns) as a tuple

3. Selection & Filtering

Selecting data is one of the most common tasks. Pandas provides multiple intuitive ways to do this.

Column Selection

# Select a single column (returns a Series)
ages = df['Age']

# Select multiple columns (returns a DataFrame)
subset = df[['Name', 'Salary']]

Row Selection using .loc and .iloc

  • .loc is label-based: references rows/columns by their row labels or column names.
  • .iloc is integer-position-based: references rows/columns by their 0-indexed positions.
# Select the first row by position
first_row = df.iloc[0]

# Select a cell by row label and column name
val = df.loc[2, 'Name'] # Charlie

Boolean Indexing (Filtering)

To filter rows based on conditions:

# Filter employees earning more than 85,000
high_earners = df[df['Salary'] > 85000]

# Combine multiple conditions using & (AND) or | (OR)
# Always wrap conditions in parentheses!
eng_seniors = df[(df['Department'] == 'Engineering') & (df['Age'] > 30)]

4. Data Cleaning

Raw data is rarely perfect. Pandas excels at handling missing values and data type conversions.

# Check for missing values
df.isna().sum()

# Drop rows with missing values
df_clean = df.dropna()

# Fill missing values with a default/placeholder
df['Salary'] = df['Salary'].fillna(df['Salary'].mean())

# Rename columns
df = df.rename(columns={'Name': 'Full Name', 'Age': 'Years'})

5. Grouping & Aggregation

To summarize data by categories, use the Split-Apply-Combine workflow via groupby.

# Calculate the average salary by department
dept_salaries = df.groupby('Department')['Salary'].mean()

# Perform multiple aggregations at once
summary = df.groupby('Department').agg({
    'Salary': ['mean', 'min', 'max'],
    'Age': 'mean'
})

🎓 Hands-on Practice Walkthrough

Let's walk through a mini-scenario. Imagine we have the following DataFrame of store transactions:

transactions = pd.DataFrame({
    'TransactionID': [101, 102, 103, 104, 105],
    'Store': ['North', 'South', 'North', 'West', 'South'],
    'Amount': [250.50, 150.00, np.nan, 300.25, 450.00],
    'ItemCount': [3, 2, 1, 5, 4]
})

Let's answer three questions:

Question 1: Fill missing amounts with the median transaction amount.

median_amount = transactions['Amount'].median() # 275.375
transactions['Amount'] = transactions['Amount'].fillna(median_amount)

Question 2: Find transactions with a total amount greater than 200.

large_tx = transactions[transactions['Amount'] > 200]

Question 3: Find the total revenue (sum of amounts) generated by each store.

store_revenue = transactions.groupby('Store')['Amount'].sum().reset_index()
print(store_revenue)
#   Store  Amount
# 0 North  525.875
# 1 South  600.00
# 2  West  300.25

⚡ Performance Tips & Best Practices

[!warning] Don't Loop Over Rows! Never write for index, row in df.iterrows(): unless absolutely necessary. Iteration is extremely slow because it disables pandas' vectorized backend.

Instead, use Vectorization:

# ❌ Slow & non-pythonic
for i in range(len(df)):
    df.loc[i, 'Tax'] = df.loc[i, 'Salary'] * 0.1

# ✅ Fast & vectorised
df['Tax'] = df['Salary'] * 0.1

[!tip] Avoid the SettingWithCopyWarning When you slice a DataFrame and then modify it, pandas warns you that you might be editing a temporary copy rather than the original source. To prevent this, use .copy() when creating a subset you intend to modify:

# ❌ Might trigger SettingWithCopyWarning
engineering = df[df['Department'] == 'Engineering']
engineering['Bonus'] = 1000

# ✅ Clean and safe
engineering = df[df['Department'] == 'Engineering'].copy()
engineering['Bonus'] = 1000