vault backup: 2026-05-30 16:02:33

This commit is contained in:
Rainyy21
2026-05-30 16:02:33 -04:00
commit 96449f8968
43 changed files with 2837 additions and 0 deletions
+237
View File
@@ -0,0 +1,237 @@
---
title: Introduction to NumPy
tags:
- python
- numpy
- data-science
- numerical-computing
category: Lesson
difficulty: Beginner to Intermediate
created: 2026-05-26
---
# Introduction to NumPy
> [!abstract] What is NumPy?
> **NumPy** (Numerical Python) is the foundational library for scientific computing in Python. It provides a high-performance multidimensional array object, tools for working with these arrays, and linear algebra, Fourier transform, and random number capabilities.
>
> ### Why use NumPy instead of standard Python lists?
> 1. **Speed:** NumPy arrays are written in C, making mathematical operations up to 100x faster than standard Python lists.
> 2. **Memory Efficiency:** NumPy arrays use contiguous blocks of memory, whereas Python lists store pointers to objects scattered across memory.
> 3. **Vectorization:** It allows performing mathematical operations on whole arrays without writing slow `for` loops.
---
## 📐 The N-Dimensional Array (`ndarray`)
The core of NumPy is the **`ndarray`** (N-dimensional array). It is a grid of values, all of the **same type** (homogenous), indexed by a tuple of non-negative integers.
```mermaid
graph TD
A[ndarray] --> B["1D Array (Vector) <br> Shape: (n,)"]
A --> C["2D Array (Matrix) <br> Shape: (m, n)"]
A --> D["3D Array (Tensor) <br> Shape: (p, m, n)"]
```
### Essential Array Attributes
Every array has attributes that describe its structure:
```python
import numpy as np
arr = np.array([[1, 2, 3], [4, 5, 6]])
print(arr.ndim) # Number of dimensions (axes) -> 2
print(arr.shape) # Tuple representing sizes in each dimension -> (2, 3)
print(arr.size) # Total number of elements -> 6
print(arr.dtype) # Data type of the elements -> int64
```
---
## 🛠️ Creating Arrays
First, import the library using the standard alias:
```python
import numpy as np
```
### 1. From Python Lists
```python
# 1D Vector
v = np.array([1, 2, 3])
# 2D Matrix
m = np.array([[1, 2], [3, 4]])
```
### 2. Built-in Placeholders
NumPy provides functions to initialize arrays with placeholders, avoiding manual creation:
```python
# Array of zeros
zeros = np.zeros((3, 4)) # 3 rows, 4 columns
# Array of ones
ones = np.ones((2, 3), dtype=np.int32)
# Range of numbers (similar to range())
range_arr = np.arange(0, 10, 2) # [0, 2, 4, 6, 8]
# Linearly spaced numbers
linspace_arr = np.linspace(0, 1, 5) # [0.0, 0.25, 0.5, 0.75, 1.0]
# Identity Matrix
eye_matrix = np.eye(3) # 3x3 identity matrix
```
### 3. Random Number Generation
```python
# Uniform random values between [0.0, 1.0)
rand_arr = np.random.rand(2, 2)
# Standard normal distribution (mean=0, std=1)
randn_arr = np.random.randn(2, 2)
# Random integers
rand_ints = np.random.randint(1, 100, size=(5,))
```
---
## ⚡ Element-wise Operations & Vectorization
In standard Python, to add two lists element-wise, you need a list comprehension or loop. In NumPy, you do it directly.
```python
x = np.array([1, 2, 3])
y = np.array([4, 5, 6])
print(x + y) # [5, 7, 9]
print(x * y) # [4, 10, 18]
print(x ** 2) # [1, 4, 9]
```
### 📡 Broadcasting
Broadcasting is a powerful mechanism that allows NumPy to perform arithmetic operations on arrays of **different shapes**. The smaller array is "broadcast" across the larger array so that they have compatible shapes.
```python
matrix = np.array([[1, 2, 3], [4, 5, 6]])
scalar = 10
# The scalar is added to every single element
print(matrix + scalar)
# [[11, 12, 13]
# [14, 15, 16]]
```
---
## 🔍 Indexing, Slicing & Masking
### Slicing 2D Arrays
Slicing follows the format `array[row_start:row_end, col_start:col_end]`.
```python
arr = np.array([
[10, 11, 12],
[20, 21, 22],
[30, 31, 32]
])
# Get row at index 1
print(arr[1, :]) # [20, 21, 22]
# Get column at index 2
print(arr[:, 2]) # [12, 22, 32]
# Slice a subgrid (top-left 2x2)
print(arr[0:2, 0:2])
# [[10, 11]
# [20, 21]]
```
### 🎭 Boolean Masking (Conditional Filtering)
You can filter arrays using conditions. NumPy returns elements where the condition resolves to `True`.
```python
data = np.array([1, 5, 8, 12, 3, 15])
# Create a boolean mask
mask = data > 5 # [False, False, True, True, False, True]
# Filter using the mask
filtered_data = data[mask] # [8, 12, 15]
```
---
## 🧮 Common Aggregations & Axis Operations
Aggregations allow you to compute statistics over entire arrays or along specific **axes**:
* `axis=0`: Down the columns (collapses rows).
* `axis=1`: Across the rows (collapses columns).
```python
arr = np.array([[1, 2], [3, 4]])
# Sum of all elements
print(np.sum(arr)) # 10
# Sum down the columns (vertical)
print(np.sum(arr, axis=0)) # [4, 6]
# Sum across the rows (horizontal)
print(np.sum(arr, axis=1)) # [3, 7]
```
---
## 🔄 Reshaping and Transposing
You can change the shape of an array without changing its data using `.reshape()` or `.T` (Transpose).
```python
flat = np.arange(1, 7) # [1, 2, 3, 4, 5, 6]
# Reshape into a 2x3 matrix
matrix = flat.reshape(2, 3)
# [[1, 2, 3]
# [4, 5, 6]]
# Transpose matrix (swap rows and columns)
transposed = matrix.T
# [[1, 4]
# [2, 5]
# [3, 6]]
```
---
## 💡 Best Practices
> [!important] Avoid Standard Loops
> Standard loops in Python are interpreted, which adds massive overhead. Vectorized operations execute in compiled C, taking advantage of CPU caches and SIMD instructions.
>
> **Example Comparison:**
> ```python
> # ❌ Extremely Slow
> values = np.random.rand(1_000_000)
> reciprocal = [1 / x for x in values]
>
> # ✅ Near Instantaneous
> reciprocal = 1 / values
> ```
> [!tip] Use In-place Operations to Save Memory
> Instead of creating a new copy, perform calculations directly on the existing array if possible using syntax like `+=`, `-=`, or `*=`.
> ```python
> a = np.ones(1000000)
> b = np.ones(1000000)
>
> # Allocates new memory
> a = a + b
>
> # Modifies 'a' in-place (saves memory allocation time)
> a += b
> ```
+241
View File
@@ -0,0 +1,241 @@
---
title: Introduction to Pandas
tags:
- python
- pandas
- data-science
- data-analysis
category: Lesson
difficulty: Beginner to Intermediate
created: 2026-05-26
---
# Introduction to Pandas
> [!abstract] What is Pandas?
> **Pandas** is a fast, powerful, flexible, and easy-to-use open-source data analysis and manipulation library built on top of the Python programming language.
>
> The name is derived from **"Panel Data"**, an econometrics term for multidimensional structured data sets. It is the foundational library for Data Science, Machine Learning, and Data Analysis in Python, acting as the bridge between raw data files (like CSVs, Excel files, or SQL databases) and numerical/modeling libraries (like NumPy, Scikit-Learn, and PyTorch).
---
## 🏗️ Core Data Structures
Pandas simplifies data manipulation by providing two primary, highly optimized data structures:
```mermaid
graph TD
A[Pandas Data Structures] --> B[Series - 1D]
A --> C[DataFrame - 2D]
B -->|Multiple columns merged| C
C -->|Single column extracted| B
```
### 1. Series (1D)
A **Series** is a one-dimensional array-like object containing an array of data and an associated array of data labels, called its **index**. Think of it as a single column in a spreadsheet.
```python
import pandas as pd
# Creating a Series
temperatures = pd.Series([22.5, 24.0, 19.5, 21.8], name="Temp")
print(temperatures)
```
### 2. DataFrame (2D)
A **DataFrame** represents a tabular, spreadsheet-like data structure containing an ordered collection of columns, each of which can be a different value type (numeric, string, boolean, etc.). It has both a row index and a column index.
| Index | Name | Age | Department |
| :--- | :--- | :--- | :--- |
| **0** | Alice | 28 | Engineering |
| **1** | Bob | 34 | Marketing |
| **2** | Charlie | 22 | HR |
---
## 🛠️ Getting Started & Creating Data
First, make sure you have pandas imported. The universal convention is to alias it as `pd`.
```python
import pandas as pd
import numpy as np
```
### Creating DataFrames Manually
You can easily create DataFrames from dictionaries or lists:
```python
data = {
'Name': ['Alice', 'Bob', 'Charlie', 'David'],
'Age': [25, 30, 35, 40],
'Salary': [70000, 80000, 120000, 90000],
'Department': ['HR', 'Engineering', 'Engineering', 'Finance']
}
df = pd.DataFrame(data)
```
---
## 🔍 Essential Operations (The "Big Five")
### 1. Reading & Writing Data
Pandas supports a wide range of file formats out of the box.
```python
# Reading data
df_csv = pd.read_csv('employees.csv')
df_excel = pd.read_excel('sales.xlsx', sheet_name='Sheet1')
df_sql = pd.read_sql('SELECT * FROM users', database_connection)
# Writing data
df.to_csv('output.csv', index=False) # index=False avoids writing the row numbers
df.to_json('output.json')
```
### 2. Inspecting Your Data
Before doing any analysis, you must understand the shape and types of your data.
```python
df.head(2) # Returns the first 2 rows
df.tail(2) # Returns the last 2 rows
df.info() # Summary of columns, non-null counts, and data types
df.describe() # Generates descriptive statistics for numerical columns
df.shape # Returns (rows, columns) as a tuple
```
### 3. Selection & Filtering
Selecting data is one of the most common tasks. Pandas provides multiple intuitive ways to do this.
#### Column Selection
```python
# Select a single column (returns a Series)
ages = df['Age']
# Select multiple columns (returns a DataFrame)
subset = df[['Name', 'Salary']]
```
#### Row Selection using `.loc` and `.iloc`
* `.loc` is **label-based**: references rows/columns by their row labels or column names.
* `.iloc` is **integer-position-based**: references rows/columns by their 0-indexed positions.
```python
# Select the first row by position
first_row = df.iloc[0]
# Select a cell by row label and column name
val = df.loc[2, 'Name'] # Charlie
```
#### Boolean Indexing (Filtering)
To filter rows based on conditions:
```python
# Filter employees earning more than 85,000
high_earners = df[df['Salary'] > 85000]
# Combine multiple conditions using & (AND) or | (OR)
# Always wrap conditions in parentheses!
eng_seniors = df[(df['Department'] == 'Engineering') & (df['Age'] > 30)]
```
### 4. Data Cleaning
Raw data is rarely perfect. Pandas excels at handling missing values and data type conversions.
```python
# Check for missing values
df.isna().sum()
# Drop rows with missing values
df_clean = df.dropna()
# Fill missing values with a default/placeholder
df['Salary'] = df['Salary'].fillna(df['Salary'].mean())
# Rename columns
df = df.rename(columns={'Name': 'Full Name', 'Age': 'Years'})
```
### 5. Grouping & Aggregation
To summarize data by categories, use the Split-Apply-Combine workflow via `groupby`.
```python
# Calculate the average salary by department
dept_salaries = df.groupby('Department')['Salary'].mean()
# Perform multiple aggregations at once
summary = df.groupby('Department').agg({
'Salary': ['mean', 'min', 'max'],
'Age': 'mean'
})
```
---
## 🎓 Hands-on Practice Walkthrough
Let's walk through a mini-scenario. Imagine we have the following DataFrame of store transactions:
```python
transactions = pd.DataFrame({
'TransactionID': [101, 102, 103, 104, 105],
'Store': ['North', 'South', 'North', 'West', 'South'],
'Amount': [250.50, 150.00, np.nan, 300.25, 450.00],
'ItemCount': [3, 2, 1, 5, 4]
})
```
Let's answer three questions:
### Question 1: Fill missing amounts with the median transaction amount.
```python
median_amount = transactions['Amount'].median() # 275.375
transactions['Amount'] = transactions['Amount'].fillna(median_amount)
```
### Question 2: Find transactions with a total amount greater than 200.
```python
large_tx = transactions[transactions['Amount'] > 200]
```
### Question 3: Find the total revenue (sum of amounts) generated by each store.
```python
store_revenue = transactions.groupby('Store')['Amount'].sum().reset_index()
print(store_revenue)
# Store Amount
# 0 North 525.875
# 1 South 600.00
# 2 West 300.25
```
---
## ⚡ Performance Tips & Best Practices
> [!warning] Don't Loop Over Rows!
> Never write `for index, row in df.iterrows():` unless absolutely necessary. Iteration is extremely slow because it disables pandas' vectorized backend.
>
> **Instead, use Vectorization:**
> ```python
> # ❌ Slow & non-pythonic
> for i in range(len(df)):
> df.loc[i, 'Tax'] = df.loc[i, 'Salary'] * 0.1
>
> # ✅ Fast & vectorised
> df['Tax'] = df['Salary'] * 0.1
> ```
> [!tip] Avoid the SettingWithCopyWarning
> When you slice a DataFrame and then modify it, pandas warns you that you might be editing a temporary copy rather than the original source.
> To prevent this, use `.copy()` when creating a subset you intend to modify:
> ```python
> # ❌ Might trigger SettingWithCopyWarning
> engineering = df[df['Department'] == 'Engineering']
> engineering['Bonus'] = 1000
>
> # ✅ Clean and safe
> engineering = df[df['Department'] == 'Engineering'].copy()
> engineering['Bonus'] = 1000
> ```
@@ -0,0 +1,130 @@
---
title: Learn Machine Learning Like a GENIUS and Not Waste Time
channel: InfiniteCodes
url: https://youtu.be/i_LwzRVP7bg
publish_date: 2024-11-14
tags:
- machine-learning
- study-methodology
- learning-strategies
- productivity
category: Video Summary
rating: 5/5
---
# Learn Machine Learning Like a GENIUS and Not Waste Time
> [!abstract] Executive Summary
> This note summarizes the video guide **"Learn Machine Learning Like a GENIUS and Not Waste Time"** by *InfiniteCodes*. The core message is that mastering machine learning (ML) isn't about memorizing complex models or chasing every new trend. Instead, it relies on building a deep intuition of the fundamentals, practicing active learning through project building, reading official documentation, and avoiding the traps of "tutorial hell" and "vibe coding" (blindly relying on LLMs).
---
## 🗺️ The ML Learning Roadmap (Order of Operations)
Beginners often rush into complex deep learning models before understanding basic concepts. The video advocates for a strict, logical **Order of Operations**:
```mermaid
graph TD
A[1. Mathematics Foundations] -->|Linear Algebra & Calculus| B[2. Exploratory Data Analysis]
B -->|Clean, Wrangling, Feature Eng| C[3. Simple Models First]
C -->|Linear/Logistic Reg, Trees| D[4. Advanced Architectures]
D -->|CNNs, RNNs, Transformers| E[5. Real-world Deployment]
```
### 1. Mathematics Foundations
You don't need a math PhD, but you must build a working intuition of:
* **Linear Algebra**: Understanding how data is represented and manipulated as matrices and vectors.
* **Calculus**: Grasping **Gradient Descent** (derivatives, partial derivatives) as the optimization engine that allows models to learn.
### 2. Exploratory Data Analysis (EDA)
Before fitting any model, you must "interview" your data:
* Use libraries like `pandas` and `numpy` to clean and structure data.
* Perform feature engineering and visualize distributions to identify patterns.
### 3. Simple Models First
Start with highly interpretable, foundational algorithms:
* **Linear Regression** & **Logistic Regression**
* **Decision Trees**
* *Why?* They are faster to train, easier to debug, and provide a baseline for more complex models.
### 4. Advanced Architectures
Only dive into deep learning (CNNs, RNNs, Transformers, etc.) when your specific project requires it and your foundations are solid.
---
## 🧠 The "Genius" Implementation Strategy
To truly master an algorithm, do not just read about it. Use this three-step implementation loop:
```mermaid
flowchart LR
Step1[1. Code from Scratch] --> Step2[2. Use Library]
Step2 --> Step3[3. Apply to Real Data]
```
1. **Code from Scratch**: Implement the core algorithm in raw Python (using only `numpy`) to understand the underlying mathematics.
2. **Use a Library**: Implement the same algorithm using `scikit-learn` to see how it is optimized and structured in production-ready libraries.
3. **Apply to Real Data**: Train both implementations on a dataset you gathered or prepared yourself (avoiding clean "toy" datasets).
> [!example] From Scratch vs. Library Example (Linear Regression)
> Below is a comparison of how you should study an algorithm.
>
> === "From Scratch (Math Intuition)"
> ```python
> import numpy as np
>
> class SimpleLinearRegression:
> def __init__(self, lr=0.01, epochs=1000):
> self.lr = lr
> self.epochs = epochs
> self.weights = None
> self.bias = None
>
> def fit(self, X, y):
> n_samples, n_features = X.shape
> self.weights = np.zeros(n_features)
> self.bias = 0
>
> # Gradient Descent loop
> for _ in range(self.epochs):
> y_predicted = np.dot(X, self.weights) + self.bias
> dw = (1 / n_samples) * np.dot(X.T, (y_predicted - y))
> db = (1 / n_samples) * np.sum(y_predicted - y)
>
> self.weights -= self.lr * dw
> self.bias -= self.lr * db
> ```
>
> === "Using a Library (Production Standard)"
> ```python
> from sklearn.linear_model import LinearRegression
>
> # Initialize and fit
> model = LinearRegression()
> model.fit(X_train, y_train)
>
> # Make predictions
> predictions = model.predict(X_test)
> ```
---
## 🚫 Key Pitfalls to Avoid
> [!danger] 1. The "3-Month Fallacy"
> Avoid courses or tutorials promising machine learning mastery in 3 months. Transitioning to a professional level takes sustained, long-term effort and continuous learning.
> [!warning] 2. Tutorial Hell & "Vibe Coding"
> * **Tutorial Hell**: Mindlessly consuming tutorials without writing code. Rule of thumb: Watch **maximum 2 tutorials** on a topic, then immediately build something yourself.
> * **Vibe Coding**: Relying entirely on LLMs (like Cursor or Copilot) to generate code for you. If you don't understand the lines of code being generated, you are building a house of cards.
> [!important] 3. Documentation First
> Build a habit of reading the **official documentation** (e.g., PyTorch, scikit-learn, Pandas) instead of asking an AI for code snippets immediately. This builds strong neural connections and developer independence.
---
## ⚡ Soft Skills & Mindset
* **Deep Work**: Dedicate uninterrupted **90 to 120-minute** blocks to study and code. Turn off notifications and focus deeply.
* **Problem Solving**: ML is about breaking complex, abstract real-world problems down into structured data steps.
* **Community & Networking**: Share your learning journey on GitHub, LinkedIn, or community Discords. Learning in public accelerates growth and opens career opportunities.
+20
View File
@@ -0,0 +1,20 @@
---
title: LLM Notes Index
created: 2026-05-26
tags:
- llm
- index
- moc
category: index
---
# LLM Notes Index
All notes under `note/LLM/`, grouped by subfolder.
```dataview
LIST rows.file.link
FROM "note/LLM"
WHERE file.name != "index"
GROUP BY file.folder
SORT file.folder ASC
```