mirror of
https://github.com/Rainyy21/framework_note.git
synced 2026-10-10 23:40:30 -04:00
vault backup: 2026-05-30 16:02:33
This commit is contained in:
@@ -0,0 +1,237 @@
|
||||
---
|
||||
title: Introduction to NumPy
|
||||
tags:
|
||||
- python
|
||||
- numpy
|
||||
- data-science
|
||||
- numerical-computing
|
||||
category: Lesson
|
||||
difficulty: Beginner to Intermediate
|
||||
created: 2026-05-26
|
||||
---
|
||||
|
||||
# Introduction to NumPy
|
||||
|
||||
> [!abstract] What is NumPy?
|
||||
> **NumPy** (Numerical Python) is the foundational library for scientific computing in Python. It provides a high-performance multidimensional array object, tools for working with these arrays, and linear algebra, Fourier transform, and random number capabilities.
|
||||
>
|
||||
> ### Why use NumPy instead of standard Python lists?
|
||||
> 1. **Speed:** NumPy arrays are written in C, making mathematical operations up to 100x faster than standard Python lists.
|
||||
> 2. **Memory Efficiency:** NumPy arrays use contiguous blocks of memory, whereas Python lists store pointers to objects scattered across memory.
|
||||
> 3. **Vectorization:** It allows performing mathematical operations on whole arrays without writing slow `for` loops.
|
||||
|
||||
---
|
||||
|
||||
## 📐 The N-Dimensional Array (`ndarray`)
|
||||
|
||||
The core of NumPy is the **`ndarray`** (N-dimensional array). It is a grid of values, all of the **same type** (homogenous), indexed by a tuple of non-negative integers.
|
||||
|
||||
```mermaid
|
||||
graph TD
|
||||
A[ndarray] --> B["1D Array (Vector) <br> Shape: (n,)"]
|
||||
A --> C["2D Array (Matrix) <br> Shape: (m, n)"]
|
||||
A --> D["3D Array (Tensor) <br> Shape: (p, m, n)"]
|
||||
```
|
||||
|
||||
### Essential Array Attributes
|
||||
Every array has attributes that describe its structure:
|
||||
|
||||
```python
|
||||
import numpy as np
|
||||
|
||||
arr = np.array([[1, 2, 3], [4, 5, 6]])
|
||||
|
||||
print(arr.ndim) # Number of dimensions (axes) -> 2
|
||||
print(arr.shape) # Tuple representing sizes in each dimension -> (2, 3)
|
||||
print(arr.size) # Total number of elements -> 6
|
||||
print(arr.dtype) # Data type of the elements -> int64
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🛠️ Creating Arrays
|
||||
|
||||
First, import the library using the standard alias:
|
||||
```python
|
||||
import numpy as np
|
||||
```
|
||||
|
||||
### 1. From Python Lists
|
||||
```python
|
||||
# 1D Vector
|
||||
v = np.array([1, 2, 3])
|
||||
|
||||
# 2D Matrix
|
||||
m = np.array([[1, 2], [3, 4]])
|
||||
```
|
||||
|
||||
### 2. Built-in Placeholders
|
||||
NumPy provides functions to initialize arrays with placeholders, avoiding manual creation:
|
||||
|
||||
```python
|
||||
# Array of zeros
|
||||
zeros = np.zeros((3, 4)) # 3 rows, 4 columns
|
||||
|
||||
# Array of ones
|
||||
ones = np.ones((2, 3), dtype=np.int32)
|
||||
|
||||
# Range of numbers (similar to range())
|
||||
range_arr = np.arange(0, 10, 2) # [0, 2, 4, 6, 8]
|
||||
|
||||
# Linearly spaced numbers
|
||||
linspace_arr = np.linspace(0, 1, 5) # [0.0, 0.25, 0.5, 0.75, 1.0]
|
||||
|
||||
# Identity Matrix
|
||||
eye_matrix = np.eye(3) # 3x3 identity matrix
|
||||
```
|
||||
|
||||
### 3. Random Number Generation
|
||||
```python
|
||||
# Uniform random values between [0.0, 1.0)
|
||||
rand_arr = np.random.rand(2, 2)
|
||||
|
||||
# Standard normal distribution (mean=0, std=1)
|
||||
randn_arr = np.random.randn(2, 2)
|
||||
|
||||
# Random integers
|
||||
rand_ints = np.random.randint(1, 100, size=(5,))
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## ⚡ Element-wise Operations & Vectorization
|
||||
|
||||
In standard Python, to add two lists element-wise, you need a list comprehension or loop. In NumPy, you do it directly.
|
||||
|
||||
```python
|
||||
x = np.array([1, 2, 3])
|
||||
y = np.array([4, 5, 6])
|
||||
|
||||
print(x + y) # [5, 7, 9]
|
||||
print(x * y) # [4, 10, 18]
|
||||
print(x ** 2) # [1, 4, 9]
|
||||
```
|
||||
|
||||
### 📡 Broadcasting
|
||||
Broadcasting is a powerful mechanism that allows NumPy to perform arithmetic operations on arrays of **different shapes**. The smaller array is "broadcast" across the larger array so that they have compatible shapes.
|
||||
|
||||
```python
|
||||
matrix = np.array([[1, 2, 3], [4, 5, 6]])
|
||||
scalar = 10
|
||||
|
||||
# The scalar is added to every single element
|
||||
print(matrix + scalar)
|
||||
# [[11, 12, 13]
|
||||
# [14, 15, 16]]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🔍 Indexing, Slicing & Masking
|
||||
|
||||
### Slicing 2D Arrays
|
||||
Slicing follows the format `array[row_start:row_end, col_start:col_end]`.
|
||||
|
||||
```python
|
||||
arr = np.array([
|
||||
[10, 11, 12],
|
||||
[20, 21, 22],
|
||||
[30, 31, 32]
|
||||
])
|
||||
|
||||
# Get row at index 1
|
||||
print(arr[1, :]) # [20, 21, 22]
|
||||
|
||||
# Get column at index 2
|
||||
print(arr[:, 2]) # [12, 22, 32]
|
||||
|
||||
# Slice a subgrid (top-left 2x2)
|
||||
print(arr[0:2, 0:2])
|
||||
# [[10, 11]
|
||||
# [20, 21]]
|
||||
```
|
||||
|
||||
### 🎭 Boolean Masking (Conditional Filtering)
|
||||
You can filter arrays using conditions. NumPy returns elements where the condition resolves to `True`.
|
||||
|
||||
```python
|
||||
data = np.array([1, 5, 8, 12, 3, 15])
|
||||
|
||||
# Create a boolean mask
|
||||
mask = data > 5 # [False, False, True, True, False, True]
|
||||
|
||||
# Filter using the mask
|
||||
filtered_data = data[mask] # [8, 12, 15]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🧮 Common Aggregations & Axis Operations
|
||||
|
||||
Aggregations allow you to compute statistics over entire arrays or along specific **axes**:
|
||||
* `axis=0`: Down the columns (collapses rows).
|
||||
* `axis=1`: Across the rows (collapses columns).
|
||||
|
||||
```python
|
||||
arr = np.array([[1, 2], [3, 4]])
|
||||
|
||||
# Sum of all elements
|
||||
print(np.sum(arr)) # 10
|
||||
|
||||
# Sum down the columns (vertical)
|
||||
print(np.sum(arr, axis=0)) # [4, 6]
|
||||
|
||||
# Sum across the rows (horizontal)
|
||||
print(np.sum(arr, axis=1)) # [3, 7]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🔄 Reshaping and Transposing
|
||||
|
||||
You can change the shape of an array without changing its data using `.reshape()` or `.T` (Transpose).
|
||||
|
||||
```python
|
||||
flat = np.arange(1, 7) # [1, 2, 3, 4, 5, 6]
|
||||
|
||||
# Reshape into a 2x3 matrix
|
||||
matrix = flat.reshape(2, 3)
|
||||
# [[1, 2, 3]
|
||||
# [4, 5, 6]]
|
||||
|
||||
# Transpose matrix (swap rows and columns)
|
||||
transposed = matrix.T
|
||||
# [[1, 4]
|
||||
# [2, 5]
|
||||
# [3, 6]]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 💡 Best Practices
|
||||
|
||||
> [!important] Avoid Standard Loops
|
||||
> Standard loops in Python are interpreted, which adds massive overhead. Vectorized operations execute in compiled C, taking advantage of CPU caches and SIMD instructions.
|
||||
>
|
||||
> **Example Comparison:**
|
||||
> ```python
|
||||
> # ❌ Extremely Slow
|
||||
> values = np.random.rand(1_000_000)
|
||||
> reciprocal = [1 / x for x in values]
|
||||
>
|
||||
> # ✅ Near Instantaneous
|
||||
> reciprocal = 1 / values
|
||||
> ```
|
||||
|
||||
> [!tip] Use In-place Operations to Save Memory
|
||||
> Instead of creating a new copy, perform calculations directly on the existing array if possible using syntax like `+=`, `-=`, or `*=`.
|
||||
> ```python
|
||||
> a = np.ones(1000000)
|
||||
> b = np.ones(1000000)
|
||||
>
|
||||
> # Allocates new memory
|
||||
> a = a + b
|
||||
>
|
||||
> # Modifies 'a' in-place (saves memory allocation time)
|
||||
> a += b
|
||||
> ```
|
||||
@@ -0,0 +1,241 @@
|
||||
---
|
||||
title: Introduction to Pandas
|
||||
tags:
|
||||
- python
|
||||
- pandas
|
||||
- data-science
|
||||
- data-analysis
|
||||
category: Lesson
|
||||
difficulty: Beginner to Intermediate
|
||||
created: 2026-05-26
|
||||
---
|
||||
|
||||
# Introduction to Pandas
|
||||
|
||||
> [!abstract] What is Pandas?
|
||||
> **Pandas** is a fast, powerful, flexible, and easy-to-use open-source data analysis and manipulation library built on top of the Python programming language.
|
||||
>
|
||||
> The name is derived from **"Panel Data"**, an econometrics term for multidimensional structured data sets. It is the foundational library for Data Science, Machine Learning, and Data Analysis in Python, acting as the bridge between raw data files (like CSVs, Excel files, or SQL databases) and numerical/modeling libraries (like NumPy, Scikit-Learn, and PyTorch).
|
||||
|
||||
---
|
||||
|
||||
## 🏗️ Core Data Structures
|
||||
|
||||
Pandas simplifies data manipulation by providing two primary, highly optimized data structures:
|
||||
|
||||
```mermaid
|
||||
graph TD
|
||||
A[Pandas Data Structures] --> B[Series - 1D]
|
||||
A --> C[DataFrame - 2D]
|
||||
B -->|Multiple columns merged| C
|
||||
C -->|Single column extracted| B
|
||||
```
|
||||
|
||||
### 1. Series (1D)
|
||||
A **Series** is a one-dimensional array-like object containing an array of data and an associated array of data labels, called its **index**. Think of it as a single column in a spreadsheet.
|
||||
|
||||
```python
|
||||
import pandas as pd
|
||||
|
||||
# Creating a Series
|
||||
temperatures = pd.Series([22.5, 24.0, 19.5, 21.8], name="Temp")
|
||||
print(temperatures)
|
||||
```
|
||||
|
||||
### 2. DataFrame (2D)
|
||||
A **DataFrame** represents a tabular, spreadsheet-like data structure containing an ordered collection of columns, each of which can be a different value type (numeric, string, boolean, etc.). It has both a row index and a column index.
|
||||
|
||||
| Index | Name | Age | Department |
|
||||
| :--- | :--- | :--- | :--- |
|
||||
| **0** | Alice | 28 | Engineering |
|
||||
| **1** | Bob | 34 | Marketing |
|
||||
| **2** | Charlie | 22 | HR |
|
||||
|
||||
---
|
||||
|
||||
## 🛠️ Getting Started & Creating Data
|
||||
|
||||
First, make sure you have pandas imported. The universal convention is to alias it as `pd`.
|
||||
|
||||
```python
|
||||
import pandas as pd
|
||||
import numpy as np
|
||||
```
|
||||
|
||||
### Creating DataFrames Manually
|
||||
You can easily create DataFrames from dictionaries or lists:
|
||||
|
||||
```python
|
||||
data = {
|
||||
'Name': ['Alice', 'Bob', 'Charlie', 'David'],
|
||||
'Age': [25, 30, 35, 40],
|
||||
'Salary': [70000, 80000, 120000, 90000],
|
||||
'Department': ['HR', 'Engineering', 'Engineering', 'Finance']
|
||||
}
|
||||
|
||||
df = pd.DataFrame(data)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🔍 Essential Operations (The "Big Five")
|
||||
|
||||
### 1. Reading & Writing Data
|
||||
Pandas supports a wide range of file formats out of the box.
|
||||
|
||||
```python
|
||||
# Reading data
|
||||
df_csv = pd.read_csv('employees.csv')
|
||||
df_excel = pd.read_excel('sales.xlsx', sheet_name='Sheet1')
|
||||
df_sql = pd.read_sql('SELECT * FROM users', database_connection)
|
||||
|
||||
# Writing data
|
||||
df.to_csv('output.csv', index=False) # index=False avoids writing the row numbers
|
||||
df.to_json('output.json')
|
||||
```
|
||||
|
||||
### 2. Inspecting Your Data
|
||||
Before doing any analysis, you must understand the shape and types of your data.
|
||||
|
||||
```python
|
||||
df.head(2) # Returns the first 2 rows
|
||||
df.tail(2) # Returns the last 2 rows
|
||||
df.info() # Summary of columns, non-null counts, and data types
|
||||
df.describe() # Generates descriptive statistics for numerical columns
|
||||
df.shape # Returns (rows, columns) as a tuple
|
||||
```
|
||||
|
||||
### 3. Selection & Filtering
|
||||
Selecting data is one of the most common tasks. Pandas provides multiple intuitive ways to do this.
|
||||
|
||||
#### Column Selection
|
||||
```python
|
||||
# Select a single column (returns a Series)
|
||||
ages = df['Age']
|
||||
|
||||
# Select multiple columns (returns a DataFrame)
|
||||
subset = df[['Name', 'Salary']]
|
||||
```
|
||||
|
||||
#### Row Selection using `.loc` and `.iloc`
|
||||
* `.loc` is **label-based**: references rows/columns by their row labels or column names.
|
||||
* `.iloc` is **integer-position-based**: references rows/columns by their 0-indexed positions.
|
||||
|
||||
```python
|
||||
# Select the first row by position
|
||||
first_row = df.iloc[0]
|
||||
|
||||
# Select a cell by row label and column name
|
||||
val = df.loc[2, 'Name'] # Charlie
|
||||
```
|
||||
|
||||
#### Boolean Indexing (Filtering)
|
||||
To filter rows based on conditions:
|
||||
|
||||
```python
|
||||
# Filter employees earning more than 85,000
|
||||
high_earners = df[df['Salary'] > 85000]
|
||||
|
||||
# Combine multiple conditions using & (AND) or | (OR)
|
||||
# Always wrap conditions in parentheses!
|
||||
eng_seniors = df[(df['Department'] == 'Engineering') & (df['Age'] > 30)]
|
||||
```
|
||||
|
||||
### 4. Data Cleaning
|
||||
Raw data is rarely perfect. Pandas excels at handling missing values and data type conversions.
|
||||
|
||||
```python
|
||||
# Check for missing values
|
||||
df.isna().sum()
|
||||
|
||||
# Drop rows with missing values
|
||||
df_clean = df.dropna()
|
||||
|
||||
# Fill missing values with a default/placeholder
|
||||
df['Salary'] = df['Salary'].fillna(df['Salary'].mean())
|
||||
|
||||
# Rename columns
|
||||
df = df.rename(columns={'Name': 'Full Name', 'Age': 'Years'})
|
||||
```
|
||||
|
||||
### 5. Grouping & Aggregation
|
||||
To summarize data by categories, use the Split-Apply-Combine workflow via `groupby`.
|
||||
|
||||
```python
|
||||
# Calculate the average salary by department
|
||||
dept_salaries = df.groupby('Department')['Salary'].mean()
|
||||
|
||||
# Perform multiple aggregations at once
|
||||
summary = df.groupby('Department').agg({
|
||||
'Salary': ['mean', 'min', 'max'],
|
||||
'Age': 'mean'
|
||||
})
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🎓 Hands-on Practice Walkthrough
|
||||
|
||||
Let's walk through a mini-scenario. Imagine we have the following DataFrame of store transactions:
|
||||
|
||||
```python
|
||||
transactions = pd.DataFrame({
|
||||
'TransactionID': [101, 102, 103, 104, 105],
|
||||
'Store': ['North', 'South', 'North', 'West', 'South'],
|
||||
'Amount': [250.50, 150.00, np.nan, 300.25, 450.00],
|
||||
'ItemCount': [3, 2, 1, 5, 4]
|
||||
})
|
||||
```
|
||||
|
||||
Let's answer three questions:
|
||||
|
||||
### Question 1: Fill missing amounts with the median transaction amount.
|
||||
```python
|
||||
median_amount = transactions['Amount'].median() # 275.375
|
||||
transactions['Amount'] = transactions['Amount'].fillna(median_amount)
|
||||
```
|
||||
|
||||
### Question 2: Find transactions with a total amount greater than 200.
|
||||
```python
|
||||
large_tx = transactions[transactions['Amount'] > 200]
|
||||
```
|
||||
|
||||
### Question 3: Find the total revenue (sum of amounts) generated by each store.
|
||||
```python
|
||||
store_revenue = transactions.groupby('Store')['Amount'].sum().reset_index()
|
||||
print(store_revenue)
|
||||
# Store Amount
|
||||
# 0 North 525.875
|
||||
# 1 South 600.00
|
||||
# 2 West 300.25
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## ⚡ Performance Tips & Best Practices
|
||||
|
||||
> [!warning] Don't Loop Over Rows!
|
||||
> Never write `for index, row in df.iterrows():` unless absolutely necessary. Iteration is extremely slow because it disables pandas' vectorized backend.
|
||||
>
|
||||
> **Instead, use Vectorization:**
|
||||
> ```python
|
||||
> # ❌ Slow & non-pythonic
|
||||
> for i in range(len(df)):
|
||||
> df.loc[i, 'Tax'] = df.loc[i, 'Salary'] * 0.1
|
||||
>
|
||||
> # ✅ Fast & vectorised
|
||||
> df['Tax'] = df['Salary'] * 0.1
|
||||
> ```
|
||||
|
||||
> [!tip] Avoid the SettingWithCopyWarning
|
||||
> When you slice a DataFrame and then modify it, pandas warns you that you might be editing a temporary copy rather than the original source.
|
||||
> To prevent this, use `.copy()` when creating a subset you intend to modify:
|
||||
> ```python
|
||||
> # ❌ Might trigger SettingWithCopyWarning
|
||||
> engineering = df[df['Department'] == 'Engineering']
|
||||
> engineering['Bonus'] = 1000
|
||||
>
|
||||
> # ✅ Clean and safe
|
||||
> engineering = df[df['Department'] == 'Engineering'].copy()
|
||||
> engineering['Bonus'] = 1000
|
||||
> ```
|
||||
@@ -0,0 +1,130 @@
|
||||
---
|
||||
title: Learn Machine Learning Like a GENIUS and Not Waste Time
|
||||
channel: InfiniteCodes
|
||||
url: https://youtu.be/i_LwzRVP7bg
|
||||
publish_date: 2024-11-14
|
||||
tags:
|
||||
- machine-learning
|
||||
- study-methodology
|
||||
- learning-strategies
|
||||
- productivity
|
||||
category: Video Summary
|
||||
rating: 5/5
|
||||
---
|
||||
|
||||
# Learn Machine Learning Like a GENIUS and Not Waste Time
|
||||
|
||||
> [!abstract] Executive Summary
|
||||
> This note summarizes the video guide **"Learn Machine Learning Like a GENIUS and Not Waste Time"** by *InfiniteCodes*. The core message is that mastering machine learning (ML) isn't about memorizing complex models or chasing every new trend. Instead, it relies on building a deep intuition of the fundamentals, practicing active learning through project building, reading official documentation, and avoiding the traps of "tutorial hell" and "vibe coding" (blindly relying on LLMs).
|
||||
|
||||
---
|
||||
|
||||
## 🗺️ The ML Learning Roadmap (Order of Operations)
|
||||
|
||||
Beginners often rush into complex deep learning models before understanding basic concepts. The video advocates for a strict, logical **Order of Operations**:
|
||||
|
||||
```mermaid
|
||||
graph TD
|
||||
A[1. Mathematics Foundations] -->|Linear Algebra & Calculus| B[2. Exploratory Data Analysis]
|
||||
B -->|Clean, Wrangling, Feature Eng| C[3. Simple Models First]
|
||||
C -->|Linear/Logistic Reg, Trees| D[4. Advanced Architectures]
|
||||
D -->|CNNs, RNNs, Transformers| E[5. Real-world Deployment]
|
||||
```
|
||||
|
||||
### 1. Mathematics Foundations
|
||||
You don't need a math PhD, but you must build a working intuition of:
|
||||
* **Linear Algebra**: Understanding how data is represented and manipulated as matrices and vectors.
|
||||
* **Calculus**: Grasping **Gradient Descent** (derivatives, partial derivatives) as the optimization engine that allows models to learn.
|
||||
|
||||
### 2. Exploratory Data Analysis (EDA)
|
||||
Before fitting any model, you must "interview" your data:
|
||||
* Use libraries like `pandas` and `numpy` to clean and structure data.
|
||||
* Perform feature engineering and visualize distributions to identify patterns.
|
||||
|
||||
### 3. Simple Models First
|
||||
Start with highly interpretable, foundational algorithms:
|
||||
* **Linear Regression** & **Logistic Regression**
|
||||
* **Decision Trees**
|
||||
* *Why?* They are faster to train, easier to debug, and provide a baseline for more complex models.
|
||||
|
||||
### 4. Advanced Architectures
|
||||
Only dive into deep learning (CNNs, RNNs, Transformers, etc.) when your specific project requires it and your foundations are solid.
|
||||
|
||||
---
|
||||
|
||||
## 🧠 The "Genius" Implementation Strategy
|
||||
|
||||
To truly master an algorithm, do not just read about it. Use this three-step implementation loop:
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Step1[1. Code from Scratch] --> Step2[2. Use Library]
|
||||
Step2 --> Step3[3. Apply to Real Data]
|
||||
```
|
||||
|
||||
1. **Code from Scratch**: Implement the core algorithm in raw Python (using only `numpy`) to understand the underlying mathematics.
|
||||
2. **Use a Library**: Implement the same algorithm using `scikit-learn` to see how it is optimized and structured in production-ready libraries.
|
||||
3. **Apply to Real Data**: Train both implementations on a dataset you gathered or prepared yourself (avoiding clean "toy" datasets).
|
||||
|
||||
> [!example] From Scratch vs. Library Example (Linear Regression)
|
||||
> Below is a comparison of how you should study an algorithm.
|
||||
>
|
||||
> === "From Scratch (Math Intuition)"
|
||||
> ```python
|
||||
> import numpy as np
|
||||
>
|
||||
> class SimpleLinearRegression:
|
||||
> def __init__(self, lr=0.01, epochs=1000):
|
||||
> self.lr = lr
|
||||
> self.epochs = epochs
|
||||
> self.weights = None
|
||||
> self.bias = None
|
||||
>
|
||||
> def fit(self, X, y):
|
||||
> n_samples, n_features = X.shape
|
||||
> self.weights = np.zeros(n_features)
|
||||
> self.bias = 0
|
||||
>
|
||||
> # Gradient Descent loop
|
||||
> for _ in range(self.epochs):
|
||||
> y_predicted = np.dot(X, self.weights) + self.bias
|
||||
> dw = (1 / n_samples) * np.dot(X.T, (y_predicted - y))
|
||||
> db = (1 / n_samples) * np.sum(y_predicted - y)
|
||||
>
|
||||
> self.weights -= self.lr * dw
|
||||
> self.bias -= self.lr * db
|
||||
> ```
|
||||
>
|
||||
> === "Using a Library (Production Standard)"
|
||||
> ```python
|
||||
> from sklearn.linear_model import LinearRegression
|
||||
>
|
||||
> # Initialize and fit
|
||||
> model = LinearRegression()
|
||||
> model.fit(X_train, y_train)
|
||||
>
|
||||
> # Make predictions
|
||||
> predictions = model.predict(X_test)
|
||||
> ```
|
||||
|
||||
---
|
||||
|
||||
## 🚫 Key Pitfalls to Avoid
|
||||
|
||||
> [!danger] 1. The "3-Month Fallacy"
|
||||
> Avoid courses or tutorials promising machine learning mastery in 3 months. Transitioning to a professional level takes sustained, long-term effort and continuous learning.
|
||||
|
||||
> [!warning] 2. Tutorial Hell & "Vibe Coding"
|
||||
> * **Tutorial Hell**: Mindlessly consuming tutorials without writing code. Rule of thumb: Watch **maximum 2 tutorials** on a topic, then immediately build something yourself.
|
||||
> * **Vibe Coding**: Relying entirely on LLMs (like Cursor or Copilot) to generate code for you. If you don't understand the lines of code being generated, you are building a house of cards.
|
||||
|
||||
> [!important] 3. Documentation First
|
||||
> Build a habit of reading the **official documentation** (e.g., PyTorch, scikit-learn, Pandas) instead of asking an AI for code snippets immediately. This builds strong neural connections and developer independence.
|
||||
|
||||
---
|
||||
|
||||
## ⚡ Soft Skills & Mindset
|
||||
|
||||
* **Deep Work**: Dedicate uninterrupted **90 to 120-minute** blocks to study and code. Turn off notifications and focus deeply.
|
||||
* **Problem Solving**: ML is about breaking complex, abstract real-world problems down into structured data steps.
|
||||
* **Community & Networking**: Share your learning journey on GitHub, LinkedIn, or community Discords. Learning in public accelerates growth and opens career opportunities.
|
||||
@@ -0,0 +1,20 @@
|
||||
---
|
||||
title: LLM Notes Index
|
||||
created: 2026-05-26
|
||||
tags:
|
||||
- llm
|
||||
- index
|
||||
- moc
|
||||
category: index
|
||||
---
|
||||
# LLM Notes Index
|
||||
|
||||
All notes under `note/LLM/`, grouped by subfolder.
|
||||
|
||||
```dataview
|
||||
LIST rows.file.link
|
||||
FROM "note/LLM"
|
||||
WHERE file.name != "index"
|
||||
GROUP BY file.folder
|
||||
SORT file.folder ASC
|
||||
```
|
||||
@@ -0,0 +1,48 @@
|
||||
---
|
||||
title: Learn Linux - The Full Course
|
||||
channel: Boot dev
|
||||
url: https://youtu.be/v392lEyM29A?si=2QnHdGS-DdjtS5Bg
|
||||
tags:
|
||||
- linux
|
||||
- course
|
||||
- devops
|
||||
- terminal
|
||||
date: 2026-05-29
|
||||
---
|
||||
|
||||
# Learn Linux - The Full Course
|
||||
|
||||
## Overview
|
||||
This comprehensive course, led by Lane from **Boot.dev**, is designed to provide developers, DevOps engineers, and IT professionals with a solid foundation in Linux and Unix-like systems. The course focuses on moving beyond basic usage to mastering the command line, navigating filesystems, and using powerful CLI tools.
|
||||
|
||||
## Key Topics
|
||||
|
||||
### Terminals and Shells (04:03)
|
||||
- Deep dive into the fundamental interface of Linux.
|
||||
- Explaining the differences between terminals and shells.
|
||||
- How they facilitate system interaction.
|
||||
|
||||
### Filesystems (20:43)
|
||||
- Navigating the Linux directory structure.
|
||||
- Managing files and understanding data organization on disk.
|
||||
|
||||
### Permissions (51:18)
|
||||
- Critical look at the Linux security model.
|
||||
- User and group management.
|
||||
- Reading and modifying file permissions.
|
||||
|
||||
### Programs (01:13:31)
|
||||
- How software is executed in a Linux environment.
|
||||
- Process management.
|
||||
- Configuration of the system `PATH`.
|
||||
|
||||
### Input/Output (01:38:33)
|
||||
- Mastery of data manipulation using standard input/output.
|
||||
- Redirection and pipes.
|
||||
- Utilities like `grep` and `find`.
|
||||
|
||||
### Packages (02:18:43)
|
||||
- Installing, updating, and managing software dependencies and packages.
|
||||
|
||||
## Conclusion
|
||||
The course transforms the command line from a source of intimidation into a powerful tool for productivity and system control.
|
||||
@@ -0,0 +1,111 @@
|
||||
---
|
||||
tags: [neovim, lazyvim, coding-tools, ide, tutorial]
|
||||
created: 2026-05-29
|
||||
status: complete
|
||||
type: lesson
|
||||
---
|
||||
|
||||
# Master Class: In-Depth Guide to LazyVim
|
||||
|
||||
LazyVim is not just a configuration; it is a **Neovim setup framework** designed to provide a high-performance, IDE-like experience while remaining modular and easy to customize. It leverages the power of `lazy.nvim` to ensure that your editor starts instantly by loading components only when they are needed.
|
||||
|
||||
---
|
||||
|
||||
## 1. The Core Philosophy
|
||||
LazyVim is built on three pillars:
|
||||
1. **Speed:** Everything is lazy-loaded. If you aren't editing a Python file, the Python LSP doesn't load.
|
||||
2. **Sane Defaults:** It comes pre-configured with industry-standard settings for UI, indentation, and search.
|
||||
3. **Modularity:** It separates its core logic from your user configuration, allowing you to update the framework without breaking your personal tweaks.
|
||||
|
||||
---
|
||||
|
||||
## 2. The Integrated Toolbox
|
||||
LazyVim integrates several powerful tools into a cohesive workflow:
|
||||
|
||||
### A. **lazy.nvim (The Engine)**
|
||||
The heart of the system. It manages plugin installation, updates, and lazy-loading.
|
||||
- **Command:** `:Lazy`
|
||||
- **Capabilities:** Check for updates, profile startup time, and manage plugin states.
|
||||
|
||||
### B. **Mason.nvim (The Tool Manager)**
|
||||
A "package manager" for your external dependencies.
|
||||
- **Command:** `:Mason`
|
||||
- **Capabilities:** Easily install and manage LSP servers, DAP (debuggers), linters, and formatters directly from within Neovim.
|
||||
|
||||
### C. **nvim-treesitter (The Parser)**
|
||||
Provides high-performance syntax highlighting and code understanding.
|
||||
- **Command:** `:TSUpdate`
|
||||
- **Capabilities:** Better highlighting, indentation, and "incremental selection" (selecting code blocks logically).
|
||||
|
||||
### D. **Telescope.nvim / fzf-lua (The Searcher)**
|
||||
A fuzzy finder that allows you to find anything in your project.
|
||||
- **Keybinds:** `<leader>ff` (files), `<leader>/` (live grep), `<leader>sk` (keymaps).
|
||||
|
||||
---
|
||||
|
||||
## 3. Essential Keybindings & Workflow
|
||||
LazyVim uses the `<Space>` key as the **Leader**. One of its best features is `which-key.nvim`, which displays a popup showing available commands whenever you press your leader key.
|
||||
|
||||
### **Navigation**
|
||||
- `<leader>e`: Toggle **Neo-tree** (File Explorer).
|
||||
- `H` / `L`: Quickly cycle through open buffers (tabs).
|
||||
- `<leader>bb`: Switch between open buffers.
|
||||
- `<leader>fT`: Open a floating terminal.
|
||||
|
||||
### **Coding & LSP**
|
||||
- `K`: Hover documentation (show what a function/variable does).
|
||||
- `gd`: Go to definition.
|
||||
- `gr`: Go to references.
|
||||
- `<leader>ca`: **Code Actions** (Fixes, imports, refactors).
|
||||
- `<leader>cr`: Rename the symbol under the cursor project-wide.
|
||||
- `[d` / `]d`: Jump to previous/next error or warning.
|
||||
|
||||
### **Git Integration**
|
||||
- `<leader>gg`: Open **LazyGit** (a full TUI for Git).
|
||||
- `<leader>gj`: Next Git hunk.
|
||||
- `<leader>gk`: Previous Git hunk.
|
||||
|
||||
---
|
||||
|
||||
## 4. Customizing Your Setup
|
||||
LazyVim's structure is designed to be clean:
|
||||
- `lua/config/options.lua`: Global Neovim settings (e.g., line numbers, tab widths).
|
||||
- `lua/config/keymaps.lua`: Your custom keyboard shortcuts.
|
||||
- `lua/plugins/`: Any `.lua` file created here is automatically loaded as a plugin configuration.
|
||||
|
||||
### **Adding a Plugin**
|
||||
Create `lua/plugins/example.lua`:
|
||||
```lua
|
||||
return {
|
||||
"username/plugin-name",
|
||||
opts = {
|
||||
-- plugin configuration goes here
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
### **Enabling "Extras"**
|
||||
LazyVim provides "Packs" for specific needs. You can enable them in `lua/config/lazy.lua`:
|
||||
```lua
|
||||
require("lazyvim.util").plugin.setup({
|
||||
spec = {
|
||||
{ "LazyVim/LazyVim", import = "lazyvim.plugins" },
|
||||
-- Enable extras like Python, Docker, or Copilot:
|
||||
{ import = "lazyvim.plugins.extras.lang.python" },
|
||||
{ import = "lazyvim.plugins.extras.ui.mini-animate" },
|
||||
{ import = "lazyvim.plugins.extras.coding.copilot" },
|
||||
{ import = "lua.plugins" },
|
||||
},
|
||||
})
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. Why LazyVim?
|
||||
Unlike building a config from scratch (which can take months to perfect) or using a "thick" distro like LunarVim (which can feel bloated), LazyVim gives you a professional-grade starting point that feels like **your** config. It provides the "glue" that makes LSP, completion, and UI tools work together seamlessly.
|
||||
|
||||
---
|
||||
**Next Steps:**
|
||||
- Run `:LazyHealth` to check your environment.
|
||||
- Install `lazygit` on your system to enable the `<leader>gg` shortcut.
|
||||
- Explore the [Official LazyVim Docs](https://www.lazyvim.org).
|
||||
@@ -0,0 +1,82 @@
|
||||
---
|
||||
tags: [neovim, treesitter, mason, telescope, coding-tools, tutorial]
|
||||
created: 2026-05-29
|
||||
status: complete
|
||||
type: lesson
|
||||
---
|
||||
|
||||
# Deep Dive: The Power Trio (Tree-sitter, Mason, & Telescope)
|
||||
|
||||
While LazyVim provides the framework, its "superpowers" come from three specific plugins that redefine how you interact with code. Understanding these tools in-depth will allow you to master your editor.
|
||||
|
||||
---
|
||||
|
||||
## 1. Tree-sitter: The Syntactic Engine
|
||||
Traditional editors use "Regex" (Regular Expressions) to highlight code. Regex is just fancy pattern matching. **Tree-sitter** is different: it is a **parser**. It builds a concrete syntax tree (CST) of your source file.
|
||||
|
||||
### Why it matters:
|
||||
- **Semantic Awareness:** It knows the difference between a variable, a function, and a type, even if they have the same name.
|
||||
- **Incremental Parsing:** It updates the tree instantly as you type, making it incredibly fast.
|
||||
- **Language-Agnostic:** One engine powers hundreds of languages.
|
||||
|
||||
### Key Modules:
|
||||
- **Highlighting:** Precise, context-aware colors.
|
||||
- **Incremental Selection:** Logical selection (select word -> select expression -> select function).
|
||||
- **Indentation:** Uses the tree structure to know exactly where a line should start.
|
||||
|
||||
### Pro Commands:
|
||||
- `:TSInstall <lang>`: Install a specific language parser.
|
||||
- `:InspectTree`: Opens a side window showing the actual tree structure of your code.
|
||||
- `:EditQuery`: Advanced tool to create custom highlights or behavior.
|
||||
|
||||
---
|
||||
|
||||
## 2. Mason.nvim: The Tool Manager
|
||||
In the past, installing an LSP (Language Server) or a Formatter was a nightmare. You had to use `npm`, `pip`, `go install`, and `cargo` manually. **Mason** abstracts all of this.
|
||||
|
||||
### The Architecture:
|
||||
- **Isolated Binaries:** Mason installs tools in `~/.local/share/nvim/mason/bin/`. It does **not** pollute your system global PATH.
|
||||
- **The "Bridge":** Mason only *downloads* the tools. Plugins like `mason-lspconfig` are needed to connect those downloads to Neovim's internal LSP client.
|
||||
|
||||
### Essential Workflow:
|
||||
1. Open `:Mason`.
|
||||
2. Find your tool (e.g., `ruff` for Python, `gopls` for Go).
|
||||
3. Press `i` to install.
|
||||
4. Use `ensure_installed` in your config to automate this across machines.
|
||||
|
||||
---
|
||||
|
||||
## 3. Telescope.nvim: The Central Nervous System
|
||||
Telescope is a highly extendable fuzzy finder. It is the primary way you find and jump to information.
|
||||
|
||||
### The "Picker" Concept:
|
||||
Everything in Telescope is a **Picker**. A picker consists of three parts:
|
||||
1. **Finder:** Where the data comes from (files, git, buffers).
|
||||
2. **Sorter:** How it ranks the results (FZF, native).
|
||||
3. **Previewer:** Shows you the content before you click.
|
||||
|
||||
### Advanced Mappings (Inside Telescope):
|
||||
- `<C-j>` / `<C-k>`: Move through results.
|
||||
- `<Tab>`: Select multiple items.
|
||||
- `<C-q>`: Send all selected items to the **Quickfix List** (powerful for mass refactoring).
|
||||
- `<C-/>`: Show all keybindings for the current picker.
|
||||
|
||||
### Must-Have Extensions:
|
||||
- **`fzf-native`:** Uses a C-port of FZF for blazing fast sorting.
|
||||
- **`ui-select`:** Makes *all* Neovim selection menus (like Code Actions) use the Telescope UI.
|
||||
- **`undo`:** A visual search for your file's history.
|
||||
|
||||
---
|
||||
|
||||
## How They Work Together
|
||||
Imagine you are editing a file:
|
||||
1. **Mason** has installed the `Pyright` LSP server.
|
||||
2. **Tree-sitter** is providing beautiful syntax highlighting and logical selection.
|
||||
3. You realize you need to find where a function is used. You trigger **Telescope** (`lsp_references`).
|
||||
4. Telescope uses the LSP (installed by **Mason**) to find the references and displays them in its UI.
|
||||
|
||||
---
|
||||
**Next Steps:**
|
||||
- Try `:InspectTree` on a complex file to see how Tree-sitter "sees" your code.
|
||||
- Run `:Mason` and check if you have any outdated tools.
|
||||
- Experiment with `live_grep` in Telescope to find text across your entire project.
|
||||
@@ -0,0 +1,20 @@
|
||||
---
|
||||
title: Python Notes Index
|
||||
created: 2026-05-22
|
||||
tags:
|
||||
- python
|
||||
- index
|
||||
- moc
|
||||
category: index
|
||||
---
|
||||
# Python Notes Index
|
||||
|
||||
All notes under `note/python/`, grouped by subfolder.
|
||||
|
||||
```dataview
|
||||
LIST rows.file.link
|
||||
FROM "note"
|
||||
WHERE file.name != "index"
|
||||
GROUP BY file.folder
|
||||
SORT file.folder ASC
|
||||
```
|
||||
@@ -0,0 +1,176 @@
|
||||
# SQLAlchemy 2.0 Tutorial (Recipe Scraper Edition)
|
||||
|
||||
This guide covers SQLAlchemy 2.0, the industry-standard SQL toolkit and Object-Relational Mapper (ORM) for Python. We'll use the **Recipe Web Scraper** project models as our primary examples.
|
||||
|
||||
---
|
||||
|
||||
## 1. What is SQLAlchemy?
|
||||
|
||||
SQLAlchemy has two main components:
|
||||
1. **Core**: A SQL abstraction layer (SQL Expression Language, Schema definitions, Engine).
|
||||
2. **ORM**: A layer on top of Core that maps Python classes to database tables.
|
||||
|
||||
In this project, we primarily use the **ORM** to treat recipes and ingredients as Python objects.
|
||||
|
||||
---
|
||||
|
||||
## 2. Defining Models (The Modern Way)
|
||||
|
||||
SQLAlchemy 2.0 introduced a type-hint-centric way to define models using `Mapped` and `mapped_column`.
|
||||
|
||||
### The Base Class
|
||||
All models inherit from a common `Base` class created from `DeclarativeBase`.
|
||||
|
||||
```python
|
||||
from sqlalchemy.orm import DeclarativeBase
|
||||
|
||||
class Base(DeclarativeBase):
|
||||
pass
|
||||
```
|
||||
|
||||
### Example: The Recipe Model
|
||||
```python
|
||||
from sqlalchemy import String, Integer, DateTime
|
||||
from sqlalchemy.orm import Mapped, mapped_column, relationship
|
||||
from datetime import datetime
|
||||
|
||||
class Recipe(Base):
|
||||
__tablename__ = "recipes" # Name of the table in the DB
|
||||
|
||||
# Primary Key
|
||||
id: Mapped[int] = mapped_column(primary_key=True)
|
||||
|
||||
# Simple Columns (SQLAlchemy infers types from Mapped[T])
|
||||
url: Mapped[str] = mapped_column(String, unique=True, index=True)
|
||||
title: Mapped[str]
|
||||
total_time: Mapped[int | None] # Optional column (nullable=True)
|
||||
|
||||
# Column with a default value
|
||||
scraped_at: Mapped[datetime] = mapped_column(DateTime, default=datetime.utcnow)
|
||||
|
||||
# Relationships (Defined in section 4)
|
||||
ingredients: Mapped[list["Ingredient"]] = relationship(back_populates="recipe")
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Engine and Session
|
||||
|
||||
### The Engine
|
||||
The **Engine** is the starting point for any SQLAlchemy application. It manages a pool of connections to the database.
|
||||
|
||||
```python
|
||||
from sqlalchemy import create_engine
|
||||
|
||||
# SQLite: The '///' means relative path to the current directory
|
||||
engine = create_engine("sqlite:///recipes.db", echo=True)
|
||||
# echo=True logs all SQL commands to the terminal (great for debugging)
|
||||
```
|
||||
|
||||
### Creating Tables
|
||||
You can tell SQLAlchemy to create all tables defined in your models:
|
||||
```python
|
||||
Base.metadata.create_all(engine)
|
||||
```
|
||||
|
||||
### The Session
|
||||
The **Session** handles the conversation with the database. Use `sessionmaker` to create a factory for sessions.
|
||||
|
||||
```python
|
||||
from sqlalchemy.orm import sessionmaker
|
||||
|
||||
SessionLocal = sessionmaker(bind=engine)
|
||||
|
||||
# Use as a context manager to ensure the connection is closed
|
||||
with SessionLocal() as session:
|
||||
# do work here
|
||||
pass
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. Relationships (1-to-Many)
|
||||
|
||||
In our project, one `Recipe` has many `Ingredients`.
|
||||
|
||||
### Foreign Key
|
||||
The "child" table (`Ingredient`) must have a column pointing to the "parent" table (`Recipe`).
|
||||
|
||||
```python
|
||||
from sqlalchemy import ForeignKey
|
||||
|
||||
class Ingredient(Base):
|
||||
__tablename__ = "ingredients"
|
||||
id: Mapped[int] = mapped_column(primary_key=True)
|
||||
|
||||
# Links to 'recipes.id'
|
||||
recipe_id: Mapped[int] = mapped_column(ForeignKey("recipes.id"))
|
||||
|
||||
text: Mapped[str]
|
||||
|
||||
# Back-reference to the parent Recipe object
|
||||
recipe: Mapped["Recipe"] = relationship(back_populates="ingredients")
|
||||
```
|
||||
|
||||
### Cascades
|
||||
`cascade="all, delete-orphan"` ensures that if you delete a Recipe, all its Ingredients are also deleted automatically.
|
||||
|
||||
---
|
||||
|
||||
## 5. CRUD Operations
|
||||
|
||||
### Create (Insert)
|
||||
```python
|
||||
with SessionLocal() as session:
|
||||
new_recipe = Recipe(title="Pasta Carbonara", url="https://example.com/pasta")
|
||||
session.add(new_recipe)
|
||||
session.commit() # Save to DB
|
||||
```
|
||||
|
||||
### Read (Select)
|
||||
```python
|
||||
from sqlalchemy import select
|
||||
|
||||
with SessionLocal() as session:
|
||||
# 1. Get by ID
|
||||
recipe = session.get(Recipe, 1)
|
||||
|
||||
# 2. Filter by column
|
||||
stmt = select(Recipe).where(Recipe.title == "Pasta Carbonara")
|
||||
result = session.execute(stmt).scalars().first()
|
||||
|
||||
# 3. Get all
|
||||
all_recipes = session.query(Recipe).all() # Older syntax, still common
|
||||
```
|
||||
|
||||
### Update
|
||||
```python
|
||||
with SessionLocal() as session:
|
||||
recipe = session.get(Recipe, 1)
|
||||
recipe.title = "Authentic Pasta Carbonara"
|
||||
session.commit()
|
||||
```
|
||||
|
||||
### Delete
|
||||
```python
|
||||
with SessionLocal() as session:
|
||||
recipe = session.get(Recipe, 1)
|
||||
session.delete(recipe)
|
||||
session.commit()
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. Common Pitfalls
|
||||
|
||||
1. **Lazy Loading**: By default, SQLAlchemy doesn't load relationships until you access them. This can cause "N+1" performance issues. Use `joinedload` to fetch everything in one query.
|
||||
2. **Session Lifecycle**: Always use a context manager (`with session:`) or close your sessions manually.
|
||||
3. **Commit vs Flush**: `session.flush()` sends changes to the DB but doesn't permanentize them. `session.commit()` makes them permanent.
|
||||
|
||||
---
|
||||
|
||||
## 7. Next Steps: Migrations with Alembic
|
||||
|
||||
As your models change (e.g., you add a `rating` column), you shouldn't just delete the DB and start over. **Alembic** is the tool used to handle database migrations.
|
||||
|
||||
Install it with: `pip install alembic`
|
||||
@@ -0,0 +1,93 @@
|
||||
---
|
||||
title: Poetry Setup and Project Initialization Guide
|
||||
created: 2026-05-22
|
||||
tags:
|
||||
- python
|
||||
- python_tool
|
||||
- poetry
|
||||
- guide
|
||||
category: python_tool
|
||||
status: reference
|
||||
up: "[[project/index]]"
|
||||
related:
|
||||
- "[[poetry_project_ideas]]"
|
||||
- "[[__init__.py explained]]"
|
||||
source:
|
||||
author:
|
||||
published:
|
||||
---
|
||||
# Poetry Setup and Project Initialization Guide
|
||||
|
||||
This guide explains how to install Poetry and start a new Python project, based on the concepts from "Introduction to Poetry - Python Dependency Management".
|
||||
|
||||
## 1. How to Install Poetry
|
||||
|
||||
While the introductory notes focus on usage, the standard way to install Poetry is via the official installer script.
|
||||
|
||||
### macOS / Linux / WSL
|
||||
Open your terminal and run:
|
||||
```bash
|
||||
curl -sSL https://install.python-poetry.org | python3 -
|
||||
```
|
||||
|
||||
### Windows (PowerShell)
|
||||
Open PowerShell and run:
|
||||
```powershell
|
||||
(Invoke-WebRequest -Uri https://install.python-poetry.org -UseBasicParsing).Content | py -
|
||||
```
|
||||
|
||||
### Verification
|
||||
After installation, restart your terminal and verify by running:
|
||||
```bash
|
||||
poetry --version
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. How to Start a Project
|
||||
|
||||
There are two main ways to start a project with Poetry:
|
||||
|
||||
### Method A: Creating a New Project (Recommended for new folders)
|
||||
To create a new project with a predefined folder structure:
|
||||
```bash
|
||||
poetry new my-project
|
||||
```
|
||||
This creates a directory named `my-project` with the following structure:
|
||||
```text
|
||||
my-project/
|
||||
├── pyproject.toml
|
||||
├── README.md
|
||||
├── my_project/
|
||||
│ └── __init__.py
|
||||
└── tests/
|
||||
└── __init__.py
|
||||
```
|
||||
|
||||
### Method B: Initializing an Existing Project
|
||||
If you already have a project folder and want to add Poetry to it:
|
||||
1. Navigate to your project directory:
|
||||
```bash
|
||||
cd my-existing-project
|
||||
```
|
||||
2. Run the interactive initialization command mentioned in the introduction:
|
||||
```bash
|
||||
poetry init
|
||||
```
|
||||
This will walk you through creating your `pyproject.toml` file interactively.
|
||||
|
||||
---
|
||||
|
||||
## 3. Key Concepts from the Introduction
|
||||
|
||||
- **`pyproject.toml`**: The single source of truth for your project configuration (replaces `requirements.txt`, `setup.py`, etc.).
|
||||
- **Deterministic Resolution**: Poetry ensures your dependencies are resolved correctly using a lockfile (`poetry.lock`).
|
||||
- **Isolation**: Poetry automatically manages virtual environments for you, ensuring your global Python installation stays clean.
|
||||
|
||||
## 4. Basic Workflow Commands
|
||||
|
||||
Once your project is started, use these commands to manage it:
|
||||
- `poetry add <package>`: Add and install a new dependency.
|
||||
- `poetry install`: Install all dependencies defined in `pyproject.toml`.
|
||||
- `poetry shell`: Activate the project's virtual environment.
|
||||
- `poetry run <command>`: Run a command inside the virtual environment without activating it.
|
||||
@@ -0,0 +1,152 @@
|
||||
---
|
||||
title: What is __init__.py?
|
||||
created: 2026-05-22
|
||||
tags:
|
||||
- python
|
||||
- python_note
|
||||
- packaging
|
||||
- reference
|
||||
category: python_note
|
||||
status: reference
|
||||
up: "[[project/index]]"
|
||||
related:
|
||||
- "[[poetry_guide]]"
|
||||
- "[[poetry_project_ideas]]"
|
||||
source:
|
||||
author:
|
||||
published:
|
||||
---
|
||||
# What is `__init__.py`?
|
||||
|
||||
`__init__.py` is the file that tells Python "this folder is a package". When you put it inside a directory, that directory becomes importable like a module.
|
||||
|
||||
## The Core Purpose
|
||||
|
||||
Without `__init__.py` (in older Python, ≤3.2), a folder was just a folder — Python could not `import` from it. With it, the folder becomes a **package** you can do this with:
|
||||
|
||||
```python
|
||||
from recipe_scraper.scraper import fetch_recipe
|
||||
```
|
||||
|
||||
Here, `recipe_scraper/` is a package because it contains `__init__.py`.
|
||||
|
||||
> Note: Since Python 3.3, "namespace packages" allow imports without `__init__.py`, but **regular packages still use it** because it gives you more control (init code, explicit exports, IDE/tooling support).
|
||||
|
||||
---
|
||||
|
||||
## What It Does in Practice
|
||||
|
||||
### 1. Marks the folder as a package
|
||||
Even an **empty** `__init__.py` is meaningful. It signals to Python:
|
||||
> "Treat this directory as something you can import from."
|
||||
|
||||
### 2. Runs initialization code
|
||||
Anything inside `__init__.py` runs **the first time** the package is imported. This is useful for:
|
||||
- Setting up logging
|
||||
- Loading config
|
||||
- Registering plugins
|
||||
|
||||
```python
|
||||
# recipe_scraper/__init__.py
|
||||
import logging
|
||||
logging.getLogger(__name__).addHandler(logging.NullHandler())
|
||||
```
|
||||
|
||||
### 3. Controls the package's public API
|
||||
You can re-export things so users don't need to know the internal file layout:
|
||||
|
||||
```python
|
||||
# recipe_scraper/__init__.py
|
||||
from .scraper import fetch_recipe
|
||||
from .parser import parse_recipe
|
||||
from .db import save_recipe
|
||||
|
||||
__all__ = ["fetch_recipe", "parse_recipe", "save_recipe"]
|
||||
```
|
||||
|
||||
Now consumers can write the short form:
|
||||
```python
|
||||
from recipe_scraper import fetch_recipe # clean
|
||||
# instead of:
|
||||
from recipe_scraper.scraper import fetch_recipe # verbose
|
||||
```
|
||||
|
||||
### 4. Defines package metadata
|
||||
A common pattern is exposing a version string:
|
||||
```python
|
||||
# recipe_scraper/__init__.py
|
||||
__version__ = "0.1.0"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## How It Fits the Recipe Scraper Project
|
||||
|
||||
Given the folder structure from [[Recipe Web Scraper - Getting Started]]:
|
||||
|
||||
```
|
||||
recipe-scraper/
|
||||
├── pyproject.toml
|
||||
└── recipe_scraper/
|
||||
├── __init__.py ← makes this a package
|
||||
├── scraper.py
|
||||
├── parser.py
|
||||
├── db.py
|
||||
└── cli.py
|
||||
```
|
||||
|
||||
The `__init__.py` here lets you:
|
||||
|
||||
1. Run `poetry run python -m recipe_scraper.cli` — only works because `recipe_scraper` is a package.
|
||||
2. Import cleanly from anywhere in the project:
|
||||
```python
|
||||
from recipe_scraper.db import Recipe
|
||||
```
|
||||
3. Optionally expose a tidy top-level API:
|
||||
```python
|
||||
# recipe_scraper/__init__.py
|
||||
from .scraper import scrape_url
|
||||
from .db import Recipe, init_db
|
||||
|
||||
__version__ = "0.1.0"
|
||||
__all__ = ["scrape_url", "Recipe", "init_db"]
|
||||
```
|
||||
So users of your library can just do:
|
||||
```python
|
||||
import recipe_scraper
|
||||
recipe_scraper.scrape_url("https://...")
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## When to Leave It Empty vs. Fill It
|
||||
|
||||
| Situation | What to put in `__init__.py` |
|
||||
| :--- | :--- |
|
||||
| Internal-only package, no public API | Empty file |
|
||||
| You want a clean import surface | Re-exports + `__all__` |
|
||||
| Library shipped to PyPI | `__version__`, re-exports, maybe logging setup |
|
||||
| One-time setup needed (config, env) | Initialization code at the top |
|
||||
|
||||
**Rule of thumb:** start with an empty `__init__.py`. Only add code when you have a concrete reason — re-exports, version, or setup. Don't put heavy logic in `__init__.py`; it runs on every import.
|
||||
|
||||
---
|
||||
|
||||
## Common Gotchas
|
||||
|
||||
- **Heavy imports slow everything down.** If `__init__.py` imports a big library (e.g., `pandas`), every `import recipe_scraper.anything` pays that cost. Keep it light.
|
||||
- **Circular imports** often start in `__init__.py`. If `__init__.py` imports from `scraper.py`, and `scraper.py` imports from the package root, you get a cycle. Use lazy imports or restructure.
|
||||
- **Tests need it too.** A `tests/` folder usually has an empty `__init__.py` so pytest can discover test modules consistently (though pytest's `rootdir` config can avoid this).
|
||||
|
||||
---
|
||||
|
||||
## TL;DR
|
||||
|
||||
- `__init__.py` = "this folder is a Python package".
|
||||
- Can be empty — just its presence matters.
|
||||
- Use it to **re-export** the public API, set `__version__`, or run small setup.
|
||||
- Keep it light: every import of the package runs it.
|
||||
|
||||
## Related
|
||||
- [[Recipe Web Scraper - Getting Started]]
|
||||
- [[poetry_guide]] — Poetry's `poetry new` creates this file automatically
|
||||
@@ -0,0 +1,177 @@
|
||||
---
|
||||
title: __init__.py vs main.py
|
||||
created: 2026-05-22
|
||||
tags:
|
||||
- python
|
||||
- python_note
|
||||
- packaging
|
||||
- reference
|
||||
category: python_note
|
||||
status: reference
|
||||
up: "[[project/index]]"
|
||||
related:
|
||||
- "[[__init__.py explained]]"
|
||||
- "[[Recipe Web Scraper - Getting Started]]"
|
||||
source:
|
||||
author:
|
||||
published:
|
||||
---
|
||||
# `__init__.py` vs `main.py`
|
||||
|
||||
Both files live inside a Python package, but they have very different jobs. This note uses the [[Recipe Web Scraper - Getting Started|recipe scraper]] project as the running example.
|
||||
|
||||
## The Recipe Scraper Layout
|
||||
|
||||
```
|
||||
recipe-scraper/
|
||||
└── src/
|
||||
└── recipe_scraper/
|
||||
├── __init__.py ← marks this as a package
|
||||
├── main.py ← CLI entry point (you run this)
|
||||
├── scraper.py
|
||||
├── models.py
|
||||
├── db.py
|
||||
└── exporters.py
|
||||
```
|
||||
|
||||
Two files, two completely different purposes.
|
||||
|
||||
---
|
||||
|
||||
## At a Glance
|
||||
|
||||
| Aspect | `__init__.py` | `main.py` |
|
||||
| :--- | :--- | :--- |
|
||||
| **Purpose** | Marks the folder as a package; sets up the package | Holds the **program's entry point** (the code you run) |
|
||||
| **Required?** | Yes (for regular packages) | No — it's a convention, name can be anything |
|
||||
| **When it runs** | Automatically, on **any** import of the package | Only when you **explicitly run it** |
|
||||
| **Typical contents** | Re-exports, `__version__`, light setup | `main()` function, CLI parsing, `if __name__ == "__main__"` |
|
||||
| **Who calls it** | Python's import system | The user (via `python -m`, `poetry run`, etc.) |
|
||||
| **Should be lightweight?** | **Yes** — runs every import | Can be heavy — runs once when invoked |
|
||||
|
||||
---
|
||||
|
||||
## `__init__.py` — The Package Setup File
|
||||
|
||||
Its job is to tell Python "this folder is a package" and optionally configure what the package looks like from the outside.
|
||||
|
||||
```python
|
||||
# src/recipe_scraper/__init__.py
|
||||
from .scraper import scrape_url
|
||||
from .db import Recipe, init_db
|
||||
|
||||
__version__ = "0.1.0"
|
||||
__all__ = ["scrape_url", "Recipe", "init_db"]
|
||||
```
|
||||
|
||||
After this, anyone using the library can write:
|
||||
```python
|
||||
from recipe_scraper import scrape_url
|
||||
scrape_url("https://example.com/recipe")
|
||||
```
|
||||
|
||||
instead of digging into `recipe_scraper.scraper`. **It never gets "run" directly — it runs implicitly whenever the package is imported.**
|
||||
|
||||
See [[__init__.py explained]] for the full picture.
|
||||
|
||||
---
|
||||
|
||||
## `main.py` — The Program Entry Point
|
||||
|
||||
Its job is to be the **code that actually executes when you launch the program**. For the recipe scraper, it's where the CLI lives:
|
||||
|
||||
```python
|
||||
# src/recipe_scraper/main.py
|
||||
import typer
|
||||
from .scraper import scrape_url
|
||||
from .db import init_db
|
||||
|
||||
app = typer.Typer()
|
||||
|
||||
@app.command()
|
||||
def scrape(url: str):
|
||||
"""Scrape a recipe URL and save it to the database."""
|
||||
init_db()
|
||||
recipe = scrape_url(url)
|
||||
print(f"Saved: {recipe.title}")
|
||||
|
||||
@app.command()
|
||||
def list_recipes():
|
||||
"""List all saved recipes."""
|
||||
...
|
||||
|
||||
def main():
|
||||
app()
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
```
|
||||
|
||||
You run it like:
|
||||
```bash
|
||||
poetry run python -m recipe_scraper.main scrape https://example.com/recipe
|
||||
```
|
||||
|
||||
Or, if `pyproject.toml` defines an entry point pointing at `main:main`, just:
|
||||
```bash
|
||||
poetry run recipe-scraper scrape https://example.com/recipe
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## The Key Mental Model
|
||||
|
||||
> `__init__.py` answers: **"What is this package?"**
|
||||
> `main.py` answers: **"What happens when you run this program?"**
|
||||
|
||||
- Importing the package → `__init__.py` runs.
|
||||
- Running the program → `main.py` runs (and it imports things, which triggers `__init__.py`).
|
||||
|
||||
So in a typical execution:
|
||||
```
|
||||
$ poetry run python -m recipe_scraper.main scrape ...
|
||||
│
|
||||
├── Python loads `recipe_scraper` package
|
||||
│ └── runs __init__.py (sets up exports, version, etc.)
|
||||
│
|
||||
└── Python runs main.py as the module
|
||||
└── parses CLI args, calls scrape_url(), saves to DB
|
||||
```
|
||||
|
||||
`__init__.py` runs **first** (because `main.py` is inside the package), but it does almost nothing visible. `main.py` is where the actual work happens.
|
||||
|
||||
---
|
||||
|
||||
## Common Pitfalls
|
||||
|
||||
### 1. Putting CLI code in `__init__.py`
|
||||
**Don't.** It will run every time *anything* imports your package, including your tests. Keep entry-point code in `main.py` (or `cli.py`).
|
||||
|
||||
### 2. Heavy imports in `__init__.py`
|
||||
If `__init__.py` imports `pandas`, `torch`, or anything slow, every import of your package pays that cost. For the recipe scraper, avoid importing `recipe-scrapers` or SQLAlchemy at the top of `__init__.py` — import them inside the modules that need them.
|
||||
|
||||
### 3. Forgetting `if __name__ == "__main__":` in `main.py`
|
||||
Without that guard, `main.py`'s code runs whenever the module is **imported**, not just when it's **executed**. Always wrap the entry-point call:
|
||||
```python
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
```
|
||||
|
||||
### 4. Confusing `main.py` with `__main__.py`
|
||||
- `main.py` — just a convention; you have to point at it explicitly (`python -m recipe_scraper.main`).
|
||||
- `__main__.py` — a **special name**: lets you run `python -m recipe_scraper` (no `.main` needed).
|
||||
|
||||
If you want `python -m recipe_scraper` to work, rename `main.py` to `__main__.py`. Many projects keep both: `__main__.py` is a one-liner that calls into `main.py`.
|
||||
|
||||
---
|
||||
|
||||
## TL;DR
|
||||
|
||||
- **`__init__.py`** = package marker + public-API shaping. Runs on every import. Keep it light.
|
||||
- **`main.py`** = program entry point. Runs when you invoke the CLI. This is where the action is.
|
||||
- They're complementary, not alternatives. In the recipe scraper, `__init__.py` exposes the library; `main.py` is the CLI you actually run.
|
||||
|
||||
## Related
|
||||
- [[__init__.py explained]]
|
||||
- [[Recipe Web Scraper - Getting Started]]
|
||||
- [[poetry_guide]]
|
||||
Reference in New Issue
Block a user