vault backup: 2026-05-30 16:02:33

This commit is contained in:
Rainyy21
2026-05-30 16:02:33 -04:00
commit 96449f8968
43 changed files with 2837 additions and 0 deletions
+20
View File
@@ -0,0 +1,20 @@
---
title: Python Notes Index
created: 2026-05-22
tags:
- python
- index
- moc
category: index
---
# Python Notes Index
All notes under `note/python/`, grouped by subfolder.
```dataview
LIST rows.file.link
FROM "note"
WHERE file.name != "index"
GROUP BY file.folder
SORT file.folder ASC
```
@@ -0,0 +1,176 @@
# SQLAlchemy 2.0 Tutorial (Recipe Scraper Edition)
This guide covers SQLAlchemy 2.0, the industry-standard SQL toolkit and Object-Relational Mapper (ORM) for Python. We'll use the **Recipe Web Scraper** project models as our primary examples.
---
## 1. What is SQLAlchemy?
SQLAlchemy has two main components:
1. **Core**: A SQL abstraction layer (SQL Expression Language, Schema definitions, Engine).
2. **ORM**: A layer on top of Core that maps Python classes to database tables.
In this project, we primarily use the **ORM** to treat recipes and ingredients as Python objects.
---
## 2. Defining Models (The Modern Way)
SQLAlchemy 2.0 introduced a type-hint-centric way to define models using `Mapped` and `mapped_column`.
### The Base Class
All models inherit from a common `Base` class created from `DeclarativeBase`.
```python
from sqlalchemy.orm import DeclarativeBase
class Base(DeclarativeBase):
pass
```
### Example: The Recipe Model
```python
from sqlalchemy import String, Integer, DateTime
from sqlalchemy.orm import Mapped, mapped_column, relationship
from datetime import datetime
class Recipe(Base):
__tablename__ = "recipes" # Name of the table in the DB
# Primary Key
id: Mapped[int] = mapped_column(primary_key=True)
# Simple Columns (SQLAlchemy infers types from Mapped[T])
url: Mapped[str] = mapped_column(String, unique=True, index=True)
title: Mapped[str]
total_time: Mapped[int | None] # Optional column (nullable=True)
# Column with a default value
scraped_at: Mapped[datetime] = mapped_column(DateTime, default=datetime.utcnow)
# Relationships (Defined in section 4)
ingredients: Mapped[list["Ingredient"]] = relationship(back_populates="recipe")
```
---
## 3. Engine and Session
### The Engine
The **Engine** is the starting point for any SQLAlchemy application. It manages a pool of connections to the database.
```python
from sqlalchemy import create_engine
# SQLite: The '///' means relative path to the current directory
engine = create_engine("sqlite:///recipes.db", echo=True)
# echo=True logs all SQL commands to the terminal (great for debugging)
```
### Creating Tables
You can tell SQLAlchemy to create all tables defined in your models:
```python
Base.metadata.create_all(engine)
```
### The Session
The **Session** handles the conversation with the database. Use `sessionmaker` to create a factory for sessions.
```python
from sqlalchemy.orm import sessionmaker
SessionLocal = sessionmaker(bind=engine)
# Use as a context manager to ensure the connection is closed
with SessionLocal() as session:
# do work here
pass
```
---
## 4. Relationships (1-to-Many)
In our project, one `Recipe` has many `Ingredients`.
### Foreign Key
The "child" table (`Ingredient`) must have a column pointing to the "parent" table (`Recipe`).
```python
from sqlalchemy import ForeignKey
class Ingredient(Base):
__tablename__ = "ingredients"
id: Mapped[int] = mapped_column(primary_key=True)
# Links to 'recipes.id'
recipe_id: Mapped[int] = mapped_column(ForeignKey("recipes.id"))
text: Mapped[str]
# Back-reference to the parent Recipe object
recipe: Mapped["Recipe"] = relationship(back_populates="ingredients")
```
### Cascades
`cascade="all, delete-orphan"` ensures that if you delete a Recipe, all its Ingredients are also deleted automatically.
---
## 5. CRUD Operations
### Create (Insert)
```python
with SessionLocal() as session:
new_recipe = Recipe(title="Pasta Carbonara", url="https://example.com/pasta")
session.add(new_recipe)
session.commit() # Save to DB
```
### Read (Select)
```python
from sqlalchemy import select
with SessionLocal() as session:
# 1. Get by ID
recipe = session.get(Recipe, 1)
# 2. Filter by column
stmt = select(Recipe).where(Recipe.title == "Pasta Carbonara")
result = session.execute(stmt).scalars().first()
# 3. Get all
all_recipes = session.query(Recipe).all() # Older syntax, still common
```
### Update
```python
with SessionLocal() as session:
recipe = session.get(Recipe, 1)
recipe.title = "Authentic Pasta Carbonara"
session.commit()
```
### Delete
```python
with SessionLocal() as session:
recipe = session.get(Recipe, 1)
session.delete(recipe)
session.commit()
```
---
## 6. Common Pitfalls
1. **Lazy Loading**: By default, SQLAlchemy doesn't load relationships until you access them. This can cause "N+1" performance issues. Use `joinedload` to fetch everything in one query.
2. **Session Lifecycle**: Always use a context manager (`with session:`) or close your sessions manually.
3. **Commit vs Flush**: `session.flush()` sends changes to the DB but doesn't permanentize them. `session.commit()` makes them permanent.
---
## 7. Next Steps: Migrations with Alembic
As your models change (e.g., you add a `rating` column), you shouldn't just delete the DB and start over. **Alembic** is the tool used to handle database migrations.
Install it with: `pip install alembic`
+93
View File
@@ -0,0 +1,93 @@
---
title: Poetry Setup and Project Initialization Guide
created: 2026-05-22
tags:
- python
- python_tool
- poetry
- guide
category: python_tool
status: reference
up: "[[project/index]]"
related:
- "[[poetry_project_ideas]]"
- "[[__init__.py explained]]"
source:
author:
published:
---
# Poetry Setup and Project Initialization Guide
This guide explains how to install Poetry and start a new Python project, based on the concepts from "Introduction to Poetry - Python Dependency Management".
## 1. How to Install Poetry
While the introductory notes focus on usage, the standard way to install Poetry is via the official installer script.
### macOS / Linux / WSL
Open your terminal and run:
```bash
curl -sSL https://install.python-poetry.org | python3 -
```
### Windows (PowerShell)
Open PowerShell and run:
```powershell
(Invoke-WebRequest -Uri https://install.python-poetry.org -UseBasicParsing).Content | py -
```
### Verification
After installation, restart your terminal and verify by running:
```bash
poetry --version
```
---
## 2. How to Start a Project
There are two main ways to start a project with Poetry:
### Method A: Creating a New Project (Recommended for new folders)
To create a new project with a predefined folder structure:
```bash
poetry new my-project
```
This creates a directory named `my-project` with the following structure:
```text
my-project/
├── pyproject.toml
├── README.md
├── my_project/
│ └── __init__.py
└── tests/
└── __init__.py
```
### Method B: Initializing an Existing Project
If you already have a project folder and want to add Poetry to it:
1. Navigate to your project directory:
```bash
cd my-existing-project
```
2. Run the interactive initialization command mentioned in the introduction:
```bash
poetry init
```
This will walk you through creating your `pyproject.toml` file interactively.
---
## 3. Key Concepts from the Introduction
- **`pyproject.toml`**: The single source of truth for your project configuration (replaces `requirements.txt`, `setup.py`, etc.).
- **Deterministic Resolution**: Poetry ensures your dependencies are resolved correctly using a lockfile (`poetry.lock`).
- **Isolation**: Poetry automatically manages virtual environments for you, ensuring your global Python installation stays clean.
## 4. Basic Workflow Commands
Once your project is started, use these commands to manage it:
- `poetry add <package>`: Add and install a new dependency.
- `poetry install`: Install all dependencies defined in `pyproject.toml`.
- `poetry shell`: Activate the project's virtual environment.
- `poetry run <command>`: Run a command inside the virtual environment without activating it.
@@ -0,0 +1,152 @@
---
title: What is __init__.py?
created: 2026-05-22
tags:
- python
- python_note
- packaging
- reference
category: python_note
status: reference
up: "[[project/index]]"
related:
- "[[poetry_guide]]"
- "[[poetry_project_ideas]]"
source:
author:
published:
---
# What is `__init__.py`?
`__init__.py` is the file that tells Python "this folder is a package". When you put it inside a directory, that directory becomes importable like a module.
## The Core Purpose
Without `__init__.py` (in older Python, ≤3.2), a folder was just a folder — Python could not `import` from it. With it, the folder becomes a **package** you can do this with:
```python
from recipe_scraper.scraper import fetch_recipe
```
Here, `recipe_scraper/` is a package because it contains `__init__.py`.
> Note: Since Python 3.3, "namespace packages" allow imports without `__init__.py`, but **regular packages still use it** because it gives you more control (init code, explicit exports, IDE/tooling support).
---
## What It Does in Practice
### 1. Marks the folder as a package
Even an **empty** `__init__.py` is meaningful. It signals to Python:
> "Treat this directory as something you can import from."
### 2. Runs initialization code
Anything inside `__init__.py` runs **the first time** the package is imported. This is useful for:
- Setting up logging
- Loading config
- Registering plugins
```python
# recipe_scraper/__init__.py
import logging
logging.getLogger(__name__).addHandler(logging.NullHandler())
```
### 3. Controls the package's public API
You can re-export things so users don't need to know the internal file layout:
```python
# recipe_scraper/__init__.py
from .scraper import fetch_recipe
from .parser import parse_recipe
from .db import save_recipe
__all__ = ["fetch_recipe", "parse_recipe", "save_recipe"]
```
Now consumers can write the short form:
```python
from recipe_scraper import fetch_recipe # clean
# instead of:
from recipe_scraper.scraper import fetch_recipe # verbose
```
### 4. Defines package metadata
A common pattern is exposing a version string:
```python
# recipe_scraper/__init__.py
__version__ = "0.1.0"
```
---
## How It Fits the Recipe Scraper Project
Given the folder structure from [[Recipe Web Scraper - Getting Started]]:
```
recipe-scraper/
├── pyproject.toml
└── recipe_scraper/
├── __init__.py ← makes this a package
├── scraper.py
├── parser.py
├── db.py
└── cli.py
```
The `__init__.py` here lets you:
1. Run `poetry run python -m recipe_scraper.cli` — only works because `recipe_scraper` is a package.
2. Import cleanly from anywhere in the project:
```python
from recipe_scraper.db import Recipe
```
3. Optionally expose a tidy top-level API:
```python
# recipe_scraper/__init__.py
from .scraper import scrape_url
from .db import Recipe, init_db
__version__ = "0.1.0"
__all__ = ["scrape_url", "Recipe", "init_db"]
```
So users of your library can just do:
```python
import recipe_scraper
recipe_scraper.scrape_url("https://...")
```
---
## When to Leave It Empty vs. Fill It
| Situation | What to put in `__init__.py` |
| :--- | :--- |
| Internal-only package, no public API | Empty file |
| You want a clean import surface | Re-exports + `__all__` |
| Library shipped to PyPI | `__version__`, re-exports, maybe logging setup |
| One-time setup needed (config, env) | Initialization code at the top |
**Rule of thumb:** start with an empty `__init__.py`. Only add code when you have a concrete reason — re-exports, version, or setup. Don't put heavy logic in `__init__.py`; it runs on every import.
---
## Common Gotchas
- **Heavy imports slow everything down.** If `__init__.py` imports a big library (e.g., `pandas`), every `import recipe_scraper.anything` pays that cost. Keep it light.
- **Circular imports** often start in `__init__.py`. If `__init__.py` imports from `scraper.py`, and `scraper.py` imports from the package root, you get a cycle. Use lazy imports or restructure.
- **Tests need it too.** A `tests/` folder usually has an empty `__init__.py` so pytest can discover test modules consistently (though pytest's `rootdir` config can avoid this).
---
## TL;DR
- `__init__.py` = "this folder is a Python package".
- Can be empty — just its presence matters.
- Use it to **re-export** the public API, set `__version__`, or run small setup.
- Keep it light: every import of the package runs it.
## Related
- [[Recipe Web Scraper - Getting Started]]
- [[poetry_guide]] — Poetry's `poetry new` creates this file automatically
@@ -0,0 +1,177 @@
---
title: __init__.py vs main.py
created: 2026-05-22
tags:
- python
- python_note
- packaging
- reference
category: python_note
status: reference
up: "[[project/index]]"
related:
- "[[__init__.py explained]]"
- "[[Recipe Web Scraper - Getting Started]]"
source:
author:
published:
---
# `__init__.py` vs `main.py`
Both files live inside a Python package, but they have very different jobs. This note uses the [[Recipe Web Scraper - Getting Started|recipe scraper]] project as the running example.
## The Recipe Scraper Layout
```
recipe-scraper/
└── src/
└── recipe_scraper/
├── __init__.py ← marks this as a package
├── main.py ← CLI entry point (you run this)
├── scraper.py
├── models.py
├── db.py
└── exporters.py
```
Two files, two completely different purposes.
---
## At a Glance
| Aspect | `__init__.py` | `main.py` |
| :--- | :--- | :--- |
| **Purpose** | Marks the folder as a package; sets up the package | Holds the **program's entry point** (the code you run) |
| **Required?** | Yes (for regular packages) | No — it's a convention, name can be anything |
| **When it runs** | Automatically, on **any** import of the package | Only when you **explicitly run it** |
| **Typical contents** | Re-exports, `__version__`, light setup | `main()` function, CLI parsing, `if __name__ == "__main__"` |
| **Who calls it** | Python's import system | The user (via `python -m`, `poetry run`, etc.) |
| **Should be lightweight?** | **Yes** — runs every import | Can be heavy — runs once when invoked |
---
## `__init__.py` — The Package Setup File
Its job is to tell Python "this folder is a package" and optionally configure what the package looks like from the outside.
```python
# src/recipe_scraper/__init__.py
from .scraper import scrape_url
from .db import Recipe, init_db
__version__ = "0.1.0"
__all__ = ["scrape_url", "Recipe", "init_db"]
```
After this, anyone using the library can write:
```python
from recipe_scraper import scrape_url
scrape_url("https://example.com/recipe")
```
instead of digging into `recipe_scraper.scraper`. **It never gets "run" directly — it runs implicitly whenever the package is imported.**
See [[__init__.py explained]] for the full picture.
---
## `main.py` — The Program Entry Point
Its job is to be the **code that actually executes when you launch the program**. For the recipe scraper, it's where the CLI lives:
```python
# src/recipe_scraper/main.py
import typer
from .scraper import scrape_url
from .db import init_db
app = typer.Typer()
@app.command()
def scrape(url: str):
"""Scrape a recipe URL and save it to the database."""
init_db()
recipe = scrape_url(url)
print(f"Saved: {recipe.title}")
@app.command()
def list_recipes():
"""List all saved recipes."""
...
def main():
app()
if __name__ == "__main__":
main()
```
You run it like:
```bash
poetry run python -m recipe_scraper.main scrape https://example.com/recipe
```
Or, if `pyproject.toml` defines an entry point pointing at `main:main`, just:
```bash
poetry run recipe-scraper scrape https://example.com/recipe
```
---
## The Key Mental Model
> `__init__.py` answers: **"What is this package?"**
> `main.py` answers: **"What happens when you run this program?"**
- Importing the package → `__init__.py` runs.
- Running the program → `main.py` runs (and it imports things, which triggers `__init__.py`).
So in a typical execution:
```
$ poetry run python -m recipe_scraper.main scrape ...
│
├── Python loads `recipe_scraper` package
│ └── runs __init__.py (sets up exports, version, etc.)
│
└── Python runs main.py as the module
└── parses CLI args, calls scrape_url(), saves to DB
```
`__init__.py` runs **first** (because `main.py` is inside the package), but it does almost nothing visible. `main.py` is where the actual work happens.
---
## Common Pitfalls
### 1. Putting CLI code in `__init__.py`
**Don't.** It will run every time *anything* imports your package, including your tests. Keep entry-point code in `main.py` (or `cli.py`).
### 2. Heavy imports in `__init__.py`
If `__init__.py` imports `pandas`, `torch`, or anything slow, every import of your package pays that cost. For the recipe scraper, avoid importing `recipe-scrapers` or SQLAlchemy at the top of `__init__.py` — import them inside the modules that need them.
### 3. Forgetting `if __name__ == "__main__":` in `main.py`
Without that guard, `main.py`'s code runs whenever the module is **imported**, not just when it's **executed**. Always wrap the entry-point call:
```python
if __name__ == "__main__":
main()
```
### 4. Confusing `main.py` with `__main__.py`
- `main.py` — just a convention; you have to point at it explicitly (`python -m recipe_scraper.main`).
- `__main__.py` — a **special name**: lets you run `python -m recipe_scraper` (no `.main` needed).
If you want `python -m recipe_scraper` to work, rename `main.py` to `__main__.py`. Many projects keep both: `__main__.py` is a one-liner that calls into `main.py`.
---
## TL;DR
- **`__init__.py`** = package marker + public-API shaping. Runs on every import. Keep it light.
- **`main.py`** = program entry point. Runs when you invoke the CLI. This is where the action is.
- They're complementary, not alternatives. In the recipe scraper, `__init__.py` exposes the library; `main.py` is the CLI you actually run.
## Related
- [[__init__.py explained]]
- [[Recipe Web Scraper - Getting Started]]
- [[poetry_guide]]