mirror of
https://github.com/Rainyy21/framework_note.git
synced 2026-10-10 22:50:29 -04:00
vault backup: 2026-05-30 16:02:33
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
|
||||
/../project/recipe web scraper/respect robots.txt check with urlib.robotparser.md
|
||||
@@ -0,0 +1,10 @@
|
||||
---
|
||||
tags:
|
||||
- coding
|
||||
---
|
||||
```dataview
|
||||
TABLE file.tags AS Tags, file.mtime AS Modified
|
||||
FROM "devop_note"
|
||||
WHERE file.name != "00_devop_note"
|
||||
SORT file.name ASC
|
||||
```
|
||||
@@ -0,0 +1,6 @@
|
||||
---
|
||||
aliases:
|
||||
up: "[[00_devop_note]]"
|
||||
tags:
|
||||
- devop
|
||||
---
|
||||
@@ -0,0 +1,5 @@
|
||||
---
|
||||
tags:
|
||||
- devop
|
||||
up: "[[00_devop_note]]"
|
||||
---
|
||||
@@ -0,0 +1,8 @@
|
||||
---
|
||||
tags:
|
||||
- devop
|
||||
up: "[[00_devop_note]]"
|
||||
---
|
||||
# What is Maven?
|
||||
maven is a build automation tool used primarily for java projects, hosted by Apache Software Foundation. Maven projects are configured using Project Object Model (POM) in a `pom.xml` file. It's used by over 70% of Java organizations, so employers actively seek people with strong Maven skills
|
||||
|
||||
@@ -0,0 +1,175 @@
|
||||
---
|
||||
aliases:
|
||||
up: "[[00_devop_note]]"
|
||||
tags:
|
||||
- devop
|
||||
---
|
||||
|
||||
# Uvicorn and Gunicorn
|
||||
|
||||
Both are Python application servers — they run your Python web app and hand requests back and forth with a web server like Nginx. But they solve **different problems**, and you'll often see them used **together**.
|
||||
|
||||
The key split: **Gunicorn is WSGI** (synchronous), **Uvicorn is ASGI** (asynchronous).
|
||||
|
||||
## Quick refresher: WSGI vs ASGI
|
||||
|
||||
| | WSGI | ASGI |
|
||||
|---|---|---|
|
||||
| Spec | PEP 3333 | "Asynchronous Server Gateway Interface" |
|
||||
| Model | Sync, one request per worker at a time | Async, many concurrent requests per worker |
|
||||
| Frameworks | Django (classic), Flask | FastAPI, Starlette, Django (async views), Sanic |
|
||||
| Supports WebSockets? | No | Yes |
|
||||
| Supports HTTP/2, SSE? | No | Yes |
|
||||
|
||||
If your app uses `async def` view functions or WebSockets, you need ASGI.
|
||||
If it's a traditional Django/Flask app with `def` views, WSGI is fine.
|
||||
|
||||
---
|
||||
|
||||
## Gunicorn ("Green Unicorn")
|
||||
|
||||
A **WSGI** server. Mature, simple, battle-tested. The de facto default for Django/Flask in production.
|
||||
|
||||
### Why people pick it
|
||||
|
||||
- **Simple config** — most options are sensible by default
|
||||
- **Pre-fork worker model** — master process forks N workers; each worker handles one request at a time
|
||||
- **Stable** — has been the standard for years
|
||||
- **Good signal handling** — graceful reloads, zero-downtime restarts
|
||||
|
||||
### Minimal usage
|
||||
|
||||
```bash
|
||||
gunicorn myproject.wsgi:application --workers 4 --bind 0.0.0.0:8000
|
||||
```
|
||||
|
||||
Or with a config file `gunicorn.conf.py`:
|
||||
|
||||
```python
|
||||
bind = "unix:/tmp/myproject.sock"
|
||||
workers = 4
|
||||
worker_class = "sync" # default
|
||||
timeout = 30
|
||||
accesslog = "-"
|
||||
errorlog = "-"
|
||||
```
|
||||
|
||||
### Worker classes
|
||||
|
||||
Gunicorn lets you swap the worker type:
|
||||
|
||||
- `sync` — default, one request at a time per worker
|
||||
- `gthread` — threaded workers (good for I/O-bound apps)
|
||||
- `gevent` / `eventlet` — async via greenlets (legacy)
|
||||
- `uvicorn.workers.UvicornWorker` — **this is the bridge to ASGI** (see below)
|
||||
|
||||
### How many workers?
|
||||
|
||||
Rule of thumb: `(2 × CPU cores) + 1`. So a 4-core box → 9 workers.
|
||||
|
||||
---
|
||||
|
||||
## Uvicorn
|
||||
|
||||
An **ASGI** server built on `uvloop` and `httptools` — both written in C, which makes it very fast. It's the standard server for FastAPI and modern async Python web apps.
|
||||
|
||||
### Why people pick it
|
||||
|
||||
- **Async-native** — handles thousands of concurrent connections per worker
|
||||
- **WebSockets + HTTP/2 support**
|
||||
- **Very fast** — uvloop is a drop-in faster replacement for asyncio's event loop
|
||||
- **Lightweight** — small dependency surface
|
||||
|
||||
### Minimal usage
|
||||
|
||||
```bash
|
||||
uvicorn myproject.main:app --host 0.0.0.0 --port 8000
|
||||
```
|
||||
|
||||
For development with auto-reload:
|
||||
|
||||
```bash
|
||||
uvicorn myproject.main:app --reload
|
||||
```
|
||||
|
||||
### Where it falls short alone
|
||||
|
||||
Uvicorn by itself is **single-process**. To use multiple CPU cores in production, you need to either:
|
||||
|
||||
1. Run multiple Uvicorn instances behind a load balancer, OR
|
||||
2. Run Uvicorn **inside Gunicorn** as worker processes (the common pattern)
|
||||
|
||||
---
|
||||
|
||||
## The common production combo: Gunicorn + Uvicorn
|
||||
|
||||
This is the standard FastAPI production setup:
|
||||
|
||||
```bash
|
||||
gunicorn myproject.main:app \
|
||||
--workers 4 \
|
||||
--worker-class uvicorn.workers.UvicornWorker \
|
||||
--bind 0.0.0.0:8000
|
||||
```
|
||||
|
||||
What's happening:
|
||||
|
||||
- **Gunicorn** is the process manager — forks workers, handles signals, restarts dead workers, manages graceful shutdowns
|
||||
- **Uvicorn** runs *inside* each Gunicorn worker — provides the ASGI event loop that runs your async code
|
||||
|
||||
You get the best of both: Gunicorn's robust process management + Uvicorn's async performance.
|
||||
|
||||
### Visual
|
||||
|
||||
```
|
||||
┌─ Uvicorn worker (async loop) ─ FastAPI app
|
||||
│
|
||||
Nginx → Gunicorn ├─ Uvicorn worker (async loop) ─ FastAPI app
|
||||
(master) │
|
||||
├─ Uvicorn worker (async loop) ─ FastAPI app
|
||||
│
|
||||
└─ Uvicorn worker (async loop) ─ FastAPI app
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Decision guide
|
||||
|
||||
| Your situation | Use |
|
||||
|---|---|
|
||||
| Django (sync), Flask | **Gunicorn** alone |
|
||||
| FastAPI, Starlette, async Django | **Gunicorn + UvicornWorker** |
|
||||
| Local dev, single FastAPI process | **Uvicorn** alone (`--reload`) |
|
||||
| WebSockets required | Must be ASGI → **Uvicorn** (alone or under Gunicorn) |
|
||||
| Legacy app, uWSGI already configured | [[uWSGI]] — no urgent need to migrate |
|
||||
|
||||
---
|
||||
|
||||
## Comparison with [[uWSGI]]
|
||||
|
||||
| | Gunicorn | Uvicorn | uWSGI |
|
||||
|---|---|---|---|
|
||||
| Protocol | WSGI | ASGI | WSGI (+ many others) |
|
||||
| Async support | No (sync workers) | Yes (native) | Limited |
|
||||
| Config complexity | Low | Low | **Very high** |
|
||||
| WebSockets | No | Yes | Partial |
|
||||
| Speed (raw) | Good | **Fastest** for async | Fast but heavy |
|
||||
| Maintained actively | Yes | Yes | Concerns |
|
||||
| Best for | Django/Flask | FastAPI | Legacy / specialized features |
|
||||
|
||||
---
|
||||
|
||||
## Common gotchas
|
||||
|
||||
1. **Don't run Uvicorn `--reload` in production** — it's a dev-only feature, has overhead and isn't safe.
|
||||
2. **Workers ≠ threads** — each Gunicorn worker is a separate Python process with its own memory. Database connections, in-memory caches, etc. are *not* shared between workers.
|
||||
3. **Timeouts matter** — Gunicorn's default `timeout=30s` will kill workers running long async tasks. Tune it for your workload.
|
||||
4. **Nginx is still recommended in front** — Uvicorn/Gunicorn don't do SSL termination, static files, or rate limiting as well as Nginx.
|
||||
5. **Logging** — by default both log to stdout/stderr; route to your log aggregator via container stdout (`-` in config).
|
||||
|
||||
---
|
||||
|
||||
## Related
|
||||
|
||||
- [[uWSGI]] — older alternative, mostly WSGI
|
||||
- [[Apache Tomcat]] — the Java equivalent (servlet container)
|
||||
@@ -0,0 +1,94 @@
|
||||
---
|
||||
aliases:
|
||||
up: "[[00_devop_note]]"
|
||||
tags:
|
||||
- devop
|
||||
---
|
||||
w
|
||||
# What is uWSGI?
|
||||
|
||||
**uWSGI** is an application server that sits between a web server (like Nginx or Apache) and a Python web application (like Django, Flask, or FastAPI). It runs your Python code and handles incoming web requests.
|
||||
|
||||
The name comes from **WSGI** (Web Server Gateway Interface), which is the standard Python spec (PEP 3333) that defines how web servers talk to Python applications. The lowercase "u" is the Greek letter μ (micro), suggesting it's lightweight — though in practice it grew into a full-featured server.
|
||||
|
||||
## Why do we need it?
|
||||
|
||||
A plain web server like Nginx doesn't know how to execute Python code. It only serves static files (HTML, CSS, images) and forwards dynamic requests elsewhere. You need a process that:
|
||||
|
||||
1. Loads your Python application into memory
|
||||
2. Receives requests from the web server
|
||||
3. Calls your Python code with the request
|
||||
4. Returns the response back to the web server
|
||||
|
||||
That middle process is uWSGI (or alternatives like Gunicorn).
|
||||
|
||||
## The typical stack
|
||||
|
||||
```
|
||||
Browser → Nginx → uWSGI → Python app (Django/Flask)
|
||||
```
|
||||
|
||||
- **Nginx**: handles SSL, static files, load balancing, gzip, caching
|
||||
- **uWSGI**: runs Python workers, manages processes/threads
|
||||
- **Python app**: your business logic
|
||||
|
||||
Nginx and uWSGI usually talk over a **Unix socket** (fast, local) or a TCP port. The protocol between them is called the **uwsgi protocol** (lowercase) — a binary protocol that's faster than plain HTTP.
|
||||
|
||||
## Key features
|
||||
|
||||
- **Process management**: spawns multiple worker processes to handle concurrent requests
|
||||
- **Threading**: each worker can run multiple threads
|
||||
- **Auto-reload**: restart workers when code changes (dev mode)
|
||||
- **Emperor mode**: one master process supervising many vassal apps
|
||||
- **Cheaper mode**: dynamically scale workers up/down based on load
|
||||
- **Multi-language**: despite the name, also supports Ruby, Perl, Go, etc.
|
||||
|
||||
## Minimal config example
|
||||
|
||||
A typical `uwsgi.ini`:
|
||||
|
||||
```ini
|
||||
[uwsgi]
|
||||
module = myproject.wsgi:application
|
||||
master = true
|
||||
processes = 4
|
||||
threads = 2
|
||||
socket = /tmp/myproject.sock
|
||||
chmod-socket = 660
|
||||
vacuum = true
|
||||
die-on-term = true
|
||||
```
|
||||
|
||||
- `module`: entry point (the WSGI callable)
|
||||
- `processes`: how many worker processes to fork
|
||||
- `socket`: where Nginx connects to
|
||||
- `vacuum`: clean up the socket on exit
|
||||
- `die-on-term`: shut down cleanly on SIGTERM
|
||||
|
||||
## Running it
|
||||
|
||||
```bash
|
||||
uwsgi --ini uwsgi.ini
|
||||
```
|
||||
|
||||
Or in production, run it under **systemd** so it restarts on failure.
|
||||
|
||||
## uWSGI vs Gunicorn
|
||||
|
||||
Both are WSGI servers. The community has largely shifted toward **Gunicorn** because:
|
||||
|
||||
- Simpler config
|
||||
- Fewer footguns
|
||||
- Easier to deploy
|
||||
|
||||
uWSGI is more feature-rich and faster in some benchmarks, but its config surface is huge (hundreds of options) and the project has had governance/maintenance concerns. For new projects, Gunicorn behind Nginx is the common default.
|
||||
|
||||
## When you'll see uWSGI
|
||||
|
||||
- Legacy Django/Flask deployments
|
||||
- Setups that need uWSGI-specific features (Emperor, cheaper, etc.)
|
||||
- Docker images for older Python web apps
|
||||
|
||||
## Related
|
||||
|
||||
- [[Apache Tomcat]] — the Java equivalent role (servlet container running Java apps behind a web server)
|
||||
@@ -0,0 +1,55 @@
|
||||
---
|
||||
tags:
|
||||
- devop
|
||||
- video
|
||||
- poetry
|
||||
- python
|
||||
up: "[[00_devop_note]]"
|
||||
related:
|
||||
- "[[Poetry]]"
|
||||
- "[[poetry_guide]]"
|
||||
- "[[poetry_project_ideas]]"
|
||||
source:
|
||||
author:
|
||||
published:
|
||||
---
|
||||
# Poetry: Dependency Management for Python
|
||||
|
||||
Poetry is a tool for **dependency management** and **packaging** in Python. It allows you to declare the libraries your project depends on and it will manage (install/update) them for you. Poetry offers a lockfile to ensure repeatable installs, and can build your project for distribution.
|
||||
|
||||
## Similarities to Maven (Java)
|
||||
|
||||
If you are familiar with Apache Maven, you can think of Poetry as serving a similar purpose in the Python ecosystem:
|
||||
|
||||
| Feature | Maven (Java) | Poetry (Python) |
|
||||
| :--- | :--- | :--- |
|
||||
| **Project Configuration** | `pom.xml` | `pyproject.toml` |
|
||||
| **Dependency Resolution** | Resolves transitive dependencies | Resolves transitive dependencies with a deterministic solver |
|
||||
| **Locking** | No direct equivalent (uses version ranges in POM) | `poetry.lock` (ensures exact versions) |
|
||||
| **Build & Packaging** | Builds JAR/WAR files | Builds Wheel and sdist packages |
|
||||
| **Publishing** | Deploys to Central/Nexus | Publishes to PyPI or private repositories |
|
||||
| **Environment Management**| Relies on external JRE/JDK | Manages Virtual Environments automatically |
|
||||
|
||||
## Key Concepts
|
||||
|
||||
### 1. `pyproject.toml`
|
||||
This is the single source of truth for your project. It replaces `setup.py`, `requirements.txt`, `setup.cfg`, `MANIFEST.in` and `pipfile`.
|
||||
|
||||
### 2. Deterministic Resolution
|
||||
Poetry comes with a custom dependency resolver that will always find a solution if one exists, or clearly explain why it failed.
|
||||
|
||||
### 3. Isolation by Default
|
||||
Poetry always runs in isolation. It either uses your existing virtual environment or creates its own to ensure that your project dependencies don't leak into your global Python installation.
|
||||
|
||||
## Basic Commands
|
||||
|
||||
- `poetry init`: Interactively create a `pyproject.toml` file.
|
||||
- `poetry add <package>`: Adds a dependency to `pyproject.toml` and installs it.
|
||||
- `poetry install`: Installs the dependencies specified in `pyproject.toml` (or `poetry.lock` if present).
|
||||
- `poetry update`: Updates dependencies to their latest versions according to `pyproject.toml` and updates the lock file.
|
||||
- `poetry run <command>`: Runs a command within the project's virtual environment.
|
||||
- `poetry shell`: Spawns a shell within the virtual environment.
|
||||
|
||||
## Why use Poetry over Pip?
|
||||
|
||||
While `pip` is the standard package installer, it doesn't handle dependency resolution or project metadata as comprehensively as Poetry. Poetry provides a more "all-in-one" experience, much like Maven does for Java, by combining dependency management, environment isolation, and packaging into a single tool.
|
||||
@@ -0,0 +1,87 @@
|
||||
---
|
||||
tags:
|
||||
- devop
|
||||
- video
|
||||
- maven
|
||||
- java
|
||||
up: "[[00_devop_note]]"
|
||||
related: "[[Maven]]"
|
||||
source: https://www.youtube.com/watch?v=T00NKLQvwYE
|
||||
author: Cameron McKenzie (TheServerSide)
|
||||
published: 2023-08-27
|
||||
---
|
||||
# Learn Apache Maven Full Tutorial in Java for Beginners
|
||||
|
||||
> An hour-long beginner-friendly course that walks through Apache Maven from scratch — installation, configuration, core commands, dependencies, plugins, and integrations with Jenkins and Docker.
|
||||
|
||||
## Overview
|
||||
Apache Maven is the most widely used build automation tool in the Java ecosystem (used by ~70% of Java organizations). This tutorial progresses from foundational concepts (installing Java/JDK and Maven) through to advanced topics like building cloud-native microservices and CI/CD integration.
|
||||
|
||||
## Topics Covered
|
||||
|
||||
### 1. Getting Started
|
||||
- What Maven is and what it enables developers to do
|
||||
- Installing the JDK (prerequisite)
|
||||
- Downloading, installing, and configuring Maven
|
||||
- Verifying the install with `mvn -v`
|
||||
|
||||
### 2. Core Maven Commands
|
||||
- `mvn compile` — compile source code
|
||||
- `mvn test` — run unit tests
|
||||
- `mvn package` — bundle compiled code into a JAR/WAR
|
||||
- `mvn install` — install artifact into local repo
|
||||
- `mvn deploy` — push artifact to a remote repo
|
||||
- `mvn clean` — wipe the `target/` directory
|
||||
|
||||
### 3. The POM File (`pom.xml`)
|
||||
- The heart of every Maven project
|
||||
- POM types (parent, aggregator, effective POM)
|
||||
- Properties and how they control build behavior
|
||||
- Project coordinates: `groupId`, `artifactId`, `version`
|
||||
|
||||
### 4. Dependency Management
|
||||
- Declaring dependencies
|
||||
- **Scopes**: `compile`, `provided`, `runtime`, `test`, `system`
|
||||
- External dependencies
|
||||
- Exclusions and optional dependencies
|
||||
- How Maven resolves transitive dependencies
|
||||
|
||||
### 5. Plugins
|
||||
- Plugins extend Maven's core functionality
|
||||
- Commonly used:
|
||||
- **Compiler Plugin** — controls the Java source/target version
|
||||
- **Surefire Plugin** — runs unit tests
|
||||
- When to add a plugin vs. rely on defaults
|
||||
|
||||
### 6. Maven vs. Gradle
|
||||
- Comparison of the two dominant JVM build tools
|
||||
- Tradeoffs: XML configuration (Maven) vs. Groovy/Kotlin DSL (Gradle)
|
||||
- Why Maven still dominates enterprise Java
|
||||
|
||||
### 7. IDE Integration
|
||||
- Creating and managing Maven projects in **IntelliJ IDEA** and **Eclipse**
|
||||
- Building projects from the IDE
|
||||
- Creating executable JARs
|
||||
- Handling multi-module projects
|
||||
|
||||
### 8. DevOps Integration
|
||||
- **Maven + Jenkins** — running builds in CI pipelines
|
||||
- **Maven + Docker** — packaging artifacts into Docker images
|
||||
- Building cloud-native microservices
|
||||
|
||||
## Key Takeaways
|
||||
- Maven is *opinionated* — follow the standard directory layout (`src/main/java`, `src/test/java`) and most things just work
|
||||
- The `pom.xml` is declarative; you describe *what* the project is, not *how* to build it
|
||||
- Dependency scopes matter — using `test` scope keeps test libraries out of production artifacts
|
||||
- Plugins are how Maven gets extended; nearly every advanced behavior is plugin-driven
|
||||
|
||||
## Actionable Next Steps
|
||||
- [ ] Install JDK and Maven; verify with `mvn -v`
|
||||
- [ ] Generate a starter project: `mvn archetype:generate`
|
||||
- [ ] Try the full lifecycle: `mvn clean package`
|
||||
- [ ] Add a dependency from [Maven Central](https://search.maven.org)
|
||||
- [ ] Wire a simple project into Jenkins or Docker
|
||||
|
||||
## Source
|
||||
- Video: [Learn Apache Maven Full Tutorial in Java for Beginners](https://www.youtube.com/watch?v=T00NKLQvwYE)
|
||||
- Companion article: [TheServerSide — Learn Maven tutorial for beginners](https://www.theserverside.com/video/Learn-Maven-tutorial-for-beginners)
|
||||
@@ -0,0 +1,32 @@
|
||||
# 🗂️ LeetCode Note Index
|
||||
|
||||
Welcome to your LeetCode problem index. This dashboard automatically aggregates and sorts all your coding notes from the `note/` directory by difficulty (**Easy**, **Medium**, and **Hard**).
|
||||
|
||||
---
|
||||
|
||||
## 📂 Difficulty Categorization
|
||||
|
||||
> [!SUCCESS] 🟢 Easy Problems
|
||||
> List of all Easy difficulty questions.
|
||||
> ```dataview
|
||||
> TABLE choice(status = "Solved", "🟢 Solved", "🔴 Unsolved") AS Status
|
||||
> FROM "leetcode/note/easy" OR #leetcode/easy
|
||||
> SORT file.name ASC
|
||||
> ```
|
||||
|
||||
> [!WARNING] 🟡 Medium Problems
|
||||
> List of all Medium difficulty questions.
|
||||
> ```dataview
|
||||
> TABLE choice(status = "Solved", "🟢 Solved", "🔴 Unsolved") AS Status
|
||||
> FROM "leetcode/note/medium" OR #leetcode/medium
|
||||
> SORT file.name ASC
|
||||
> ```
|
||||
|
||||
> [!DANGER] 🔴 Hard Problems
|
||||
> List of all Hard difficulty questions.
|
||||
> ```dataview
|
||||
> TABLE choice(status = "Solved", "🟢 Solved", "🔴 Unsolved") AS Status
|
||||
> FROM "leetcode/note/hard" OR #leetcode/hard
|
||||
> SORT file.name ASC
|
||||
> ```
|
||||
|
||||
@@ -0,0 +1,60 @@
|
||||
---
|
||||
id: 1346
|
||||
title: Check If N and Its Double Exist
|
||||
difficulty: Easy
|
||||
tags:
|
||||
- array
|
||||
- hash-table
|
||||
- two-pointers
|
||||
- binary-search
|
||||
- sorting
|
||||
status: Solved
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/check-if-n-and-its-double-exist/
|
||||
review_needed: false
|
||||
---
|
||||
|
||||
# 1346. Check If N and Its Double Exist
|
||||
|
||||
> [!info] **Problem Link**: [LeetCode - Check If N and Its Double Exist](https://leetcode.com/problems/check-if-n-and-its-double-exist/)
|
||||
|
||||
## 📝 Problem Description
|
||||
|
||||
Given an array `arr` of integers, check if there exist two indices `i` and `j` such that :
|
||||
|
||||
- ` i != j`
|
||||
- `0 <= i, j < arr.length`
|
||||
- `arr[i] == 2 * arr[j]`
|
||||
|
||||
|
||||
---
|
||||
|
||||
### 📥 Example 1
|
||||
> **Input:** `arr = [10,2,5,3]`
|
||||
> **Output:** `true`
|
||||
> **Explanation:** For `i = 0` and `j = 2`, `arr[i] == 10 == 2 * 5 == 2 * arr[j]`
|
||||
|
||||
### 📥 Example 2
|
||||
> **Input:** `arr = [3,1,7,11]`
|
||||
> **Output:** `false`
|
||||
> **Explanation:** There is no i and j that satisfy the conditions.
|
||||
|
||||
|
||||
---
|
||||
|
||||
## 💡 Approaches & Explanations
|
||||
Have a mem that hold int that you have see before. for each int in the array you check if you seen double or half in the mem if it is than return True. In the end return False
|
||||
|
||||
## 💻 Code Implementations
|
||||
|
||||
### Python3
|
||||
```python
|
||||
class Solution:
|
||||
def checkIfExist(self, arr: List[int]) -> bool:
|
||||
mem = []
|
||||
for idx , i in enumerate(arr):
|
||||
if i*2 in mem or i/2 in mem:
|
||||
return True
|
||||
mem.append(i)
|
||||
return False
|
||||
```
|
||||
@@ -0,0 +1,58 @@
|
||||
---
|
||||
id: 169
|
||||
title: Majority Element
|
||||
difficulty: Easy
|
||||
tags:
|
||||
- array
|
||||
- hash-table
|
||||
- sorting
|
||||
- counting
|
||||
status: Solved
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/majority-element/
|
||||
review_needed: false
|
||||
---
|
||||
|
||||
# 169. Majority Element
|
||||
|
||||
> [!info] **Problem Link**: [LeetCode - Majority Element](https://leetcode.com/problems/majority-element/)
|
||||
|
||||
## 📝 Problem Description
|
||||
|
||||
Given an array `nums` of size `n`, return the majority element.
|
||||
|
||||
The majority element is the element that appears more than `[n / 2]` times. You may assume that the majority element always exists in the array.
|
||||
|
||||
|
||||
---
|
||||
|
||||
### 📥 Example 1
|
||||
> **Input:** `nums = [3,2,3]`
|
||||
> **Output:** `3`
|
||||
|
||||
### 📥 Example 2
|
||||
> **Input:** `nums = [2,2,1,1,1,2,2]`
|
||||
> **Output:** `2`
|
||||
|
||||
|
||||
---
|
||||
|
||||
## 💡 Approaches & Explanations
|
||||
have cont, if cont is equal to 0 the res become the highest amount
|
||||
if i equal to the res than add one to count anything else cont -1
|
||||
return res at the end
|
||||
|
||||
## 💻 Code Implementations
|
||||
|
||||
### Python3
|
||||
```python
|
||||
class Solution:
|
||||
def majorityElement(self, nums: List[int]) -> int:
|
||||
cont = 0
|
||||
res = None
|
||||
for i in nums:
|
||||
if cont == 0:
|
||||
res = i
|
||||
cont +=1 if res == i else -1
|
||||
return res
|
||||
```
|
||||
@@ -0,0 +1,82 @@
|
||||
---
|
||||
id: 1
|
||||
title: Two Sum
|
||||
difficulty: Easy
|
||||
tags:
|
||||
- array
|
||||
- hash-table
|
||||
status: Solved
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/two-sum/
|
||||
review_needed: false
|
||||
---
|
||||
|
||||
# 1. Two Sum
|
||||
|
||||
> [!info] **Problem Link**: [LeetCode - Two Sum](https://leetcode.com/problems/two-sum/)
|
||||
|
||||
## 📝 Problem Description
|
||||
|
||||
Given an array of integers `nums` and an integer `target`, return *indices of the two numbers such that they add up to `target`*.
|
||||
|
||||
You may assume that each input would have ***exactly* one solution**, and you may not use the *same* element twice.
|
||||
|
||||
You can return the answer in any order.
|
||||
|
||||
---
|
||||
|
||||
### 📥 Example 1
|
||||
> **Input:** `nums = [2,7,11,15]`, `target = 9`
|
||||
> **Output:** `[0,1]`
|
||||
> **Explanation:** Because `nums[0] + nums[1] == 9`, we return `[0, 1]`.
|
||||
|
||||
### 📥 Example 2
|
||||
> **Input:** `nums = [3,2,4]`, `target = 6`
|
||||
> **Output:** `[1,2]`
|
||||
|
||||
### 📥 Example 3
|
||||
> **Input:** `nums = [3,3]`, `target = 6`
|
||||
> **Output:** `[0,1]`
|
||||
|
||||
---
|
||||
|
||||
## 💡 Approaches & Explanations
|
||||
|
||||
### Approach 1: Hash Map (One-Pass) — *Optimal*
|
||||
The optimal approach is to use a hash map to keep track of the numbers we have seen so far and their indices. As we iterate through the array, we check if the complement (`target - nums[i]`) already exists in our hash map.
|
||||
- If it does, we found the pair and return their indices.
|
||||
- If it doesn't, we add the current number and its index to the hash map.
|
||||
|
||||
#### 📊 Complexity Analysis
|
||||
- **Time Complexity:** $\mathcal{O}(N)$ where $N$ is the number of elements in the array. We traverse the list containing $N$ elements only once, and lookup in the hash table takes $\mathcal{O}(1)$ time.
|
||||
- **Space Complexity:** $\mathcal{O}(N)$ since we store at most $N$ elements in the hash map.
|
||||
|
||||
---
|
||||
|
||||
### Approach 2: Brute Force
|
||||
Compare every pair of numbers to see if their sum equals the target.
|
||||
- **Time Complexity:** $\mathcal{O}(N^2)$
|
||||
- **Space Complexity:** $\mathcal{O}(1)$
|
||||
|
||||
---
|
||||
|
||||
## 💻 Code Implementations
|
||||
|
||||
### Python3
|
||||
```python
|
||||
class Solution:
|
||||
def twoSum(self, nums: List[int], target: int) -> List[int]:
|
||||
seen = {} # val -> index
|
||||
for i, num in enumerate(nums):
|
||||
complement = target - num
|
||||
if complement in seen:
|
||||
return [seen[complement], i]
|
||||
seen[num] = i
|
||||
return []
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🧠 Key Takeaways & Lessons
|
||||
- **The Complement Trick:** When looking for a pair that sums to a target, rephrase the search: instead of looking for $A + B = \text{target}$, look for $\text{complement} = \text{target} - A$ that is already stored.
|
||||
- **Hash Map for $\mathcal{O}(1)$ Lookups:** Trading memory (space complexity) for time complexity is a common pattern in array search problems.
|
||||
@@ -0,0 +1,88 @@
|
||||
---
|
||||
id: 217
|
||||
title: Contains Duplicate
|
||||
difficulty: Easy
|
||||
tags:
|
||||
- array
|
||||
- hash-table
|
||||
status: Solved
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/contains-duplicate/
|
||||
review_needed: false
|
||||
---
|
||||
|
||||
# 217. Contains Duplicate
|
||||
|
||||
> [!info] **Problem Link**: [LeetCode - Contains Duplicate](https://leetcode.com/problems/contains-duplicate/)
|
||||
|
||||
## 📝 Problem Description
|
||||
Given an integer array `nums`, return `true` if *any value appears at least twice in the array*, and return `false` if every element is distinct.
|
||||
|
||||
---
|
||||
|
||||
### 📥 Example 1
|
||||
> **Input:** `nums = [1,2,3,1]`
|
||||
> **Output:** `true`
|
||||
> **Explanation:** The element 1 occurs at the indices 0 and 3.
|
||||
|
||||
### 📥 Example 2
|
||||
> **Input:** `nums = [1,2,3,4]`
|
||||
> **Output:** `false`
|
||||
|
||||
### 📥 Example 3
|
||||
> **Input:** `nums = [1,1,1,3,3,4,3,2,4,2]`
|
||||
> **Output:** `true`
|
||||
|
||||
---
|
||||
|
||||
## 💡 Approaches & Explanations
|
||||
|
||||
### Approach 1: Hash Set (Length Comparison)
|
||||
The simplest way to check for duplicates in Python is to convert the array `nums` into a set. A set only contains unique elements, so:
|
||||
- If there are duplicates, the length of the set will be less than the length of the array.
|
||||
- If all elements are unique, the lengths will be equal.
|
||||
|
||||
#### 📊 Complexity Analysis
|
||||
- **Time Complexity:** $\mathcal{O}(N)$ where $N$ is the number of elements in the array. Converting an array to a set requires traversing the entire array and inserting each element.
|
||||
- **Space Complexity:** $\mathcal{O}(N)$ as we store up to $N$ unique elements in the set.
|
||||
|
||||
---
|
||||
|
||||
### Approach 2: Hash Set (Early Return / One-Pass) — *Alternative*
|
||||
Instead of converting the entire array to a set, we can iterate through the array and store elements in a set as we go. If we encounter an element that is already in the set, we can return `true` immediately. This avoids processing the rest of the array.
|
||||
|
||||
#### 📊 Complexity Analysis
|
||||
- **Time Complexity:** $\mathcal{O}(N)$ in the worst case (no duplicates). In the best case, it can be $\mathcal{O}(1)$ if a duplicate is found at the beginning.
|
||||
- **Space Complexity:** $\mathcal{O}(N)$ to store the visited elements.
|
||||
|
||||
---
|
||||
|
||||
## 💻 Code Implementations
|
||||
|
||||
### Python3
|
||||
|
||||
#### Option A: Length Comparison (Concise)
|
||||
```python
|
||||
class Solution:
|
||||
def containsDuplicate(self, nums: List[int]) -> bool:
|
||||
return len(nums) != len(set(nums))
|
||||
```
|
||||
|
||||
#### Option B: Early Return (Optimal for large lists with early duplicates)
|
||||
```python
|
||||
class Solution:
|
||||
def containsDuplicate(self, nums: List[int]) -> bool:
|
||||
seen = set()
|
||||
for num in nums:
|
||||
if num in seen:
|
||||
return True
|
||||
seen.add(num)
|
||||
return False
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🧠 Key Takeaways & Lessons
|
||||
- **Hash Set for Uniqueness:** Sets are the go-to data structure when you need to verify uniqueness or look up elements in $\mathcal{O}(1)$ time.
|
||||
- **Early Return Optimization:** While converting the whole list to a set is clean and concise, iterating and returning early when a duplicate is found can save time and memory in practice.
|
||||
- **Time-Space Trade-off:** We use extra space ($\mathcal{O}(N)$ memory) to achieve linear time complexity ($\mathcal{O}(N)$) instead of a brute-force search ($\mathcal{O}(N^2)$).
|
||||
@@ -0,0 +1,13 @@
|
||||
---
|
||||
id: 242
|
||||
title: Valid Anagram
|
||||
difficulty: Easy
|
||||
tags:
|
||||
- hash-table
|
||||
- string
|
||||
- sorting
|
||||
status: unsolve
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/valid-anagram/
|
||||
review_needed: false
|
||||
---
|
||||
@@ -0,0 +1,12 @@
|
||||
---
|
||||
id: 290
|
||||
title: Word Pattern
|
||||
difficulty: Easy
|
||||
tags:
|
||||
- hash-table
|
||||
- string
|
||||
status: unsolve
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/word-pattern/
|
||||
review_needed: false
|
||||
---
|
||||
@@ -0,0 +1,12 @@
|
||||
---
|
||||
id: 724
|
||||
title: Find Pivot Index
|
||||
difficulty: Easy
|
||||
tags:
|
||||
- array
|
||||
- prefix-sum
|
||||
status: unsolve
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/find-pivot-index/
|
||||
review_needed: false
|
||||
---
|
||||
@@ -0,0 +1,13 @@
|
||||
---
|
||||
id: 30
|
||||
title: Substring with Concatenation of All Words
|
||||
difficulty: Hard
|
||||
tags:
|
||||
- hash-table
|
||||
- string
|
||||
- sliding-window
|
||||
status: unsolve
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/substring-with-concatenation-of-all-words/
|
||||
review_needed: false
|
||||
---
|
||||
@@ -0,0 +1,15 @@
|
||||
---
|
||||
id: 381
|
||||
title: Insert Delete GetRandom O(1) - Duplicates allowed
|
||||
difficulty: Hard
|
||||
tags:
|
||||
- array
|
||||
- hash-table
|
||||
- math
|
||||
- randomized
|
||||
- design
|
||||
status: unsolve
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/insert-delete-getrandom-o1-duplicates-allowed/
|
||||
review_needed: false
|
||||
---
|
||||
@@ -0,0 +1,12 @@
|
||||
---
|
||||
id: 41
|
||||
title: First Missing Positive
|
||||
difficulty: Hard
|
||||
tags:
|
||||
- array
|
||||
- hash-table
|
||||
status: unsolve
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/first-missing-positive/
|
||||
review_needed: false
|
||||
---
|
||||
@@ -0,0 +1,13 @@
|
||||
---
|
||||
id: 128
|
||||
title: Longest Consecutive Sequence
|
||||
difficulty: Medium
|
||||
tags:
|
||||
- array
|
||||
- hash-table
|
||||
- union-find
|
||||
status: unsolve
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/longest-consecutive-sequence/
|
||||
review_needed: false
|
||||
---
|
||||
@@ -0,0 +1,13 @@
|
||||
---
|
||||
id: 2017
|
||||
title: Grid Game
|
||||
difficulty: Medium
|
||||
tags:
|
||||
- array
|
||||
- matrix
|
||||
- prefix-sum
|
||||
status: unsolve
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/grid-game/
|
||||
review_needed: false
|
||||
---
|
||||
@@ -0,0 +1,12 @@
|
||||
---
|
||||
id: 238
|
||||
title: Product of Array Except Self
|
||||
difficulty: Medium
|
||||
tags:
|
||||
- array
|
||||
- prefix-sum
|
||||
status: unsolve
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/product-of-array-except-self/
|
||||
review_needed: false
|
||||
---
|
||||
@@ -0,0 +1,14 @@
|
||||
---
|
||||
id: 271
|
||||
title: Encode and Decode Strings
|
||||
difficulty: Medium
|
||||
tags:
|
||||
- array
|
||||
- hash-table
|
||||
- string
|
||||
- design
|
||||
status: unsolve
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/encode-and-decode-strings/
|
||||
review_needed: false
|
||||
---
|
||||
@@ -0,0 +1,17 @@
|
||||
---
|
||||
id: 347
|
||||
title: Top K Frequent Elements
|
||||
difficulty: Medium
|
||||
tags:
|
||||
- array
|
||||
- hash-table
|
||||
- divide-and-conquer
|
||||
- sorting
|
||||
- heap-priority-queue
|
||||
- bucket-sort
|
||||
- quickselect
|
||||
status: unsolve
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/top-k-frequent-elements/
|
||||
review_needed: false
|
||||
---
|
||||
@@ -0,0 +1,13 @@
|
||||
---
|
||||
id: 36
|
||||
title: Valid Sudoku
|
||||
difficulty: Medium
|
||||
tags:
|
||||
- array
|
||||
- hash-table
|
||||
- matrix
|
||||
status: unsolve
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/valid-sudoku/
|
||||
review_needed: false
|
||||
---
|
||||
@@ -0,0 +1,15 @@
|
||||
---
|
||||
id: 380
|
||||
title: Insert Delete GetRandom O(1)
|
||||
difficulty: Medium
|
||||
tags:
|
||||
- array
|
||||
- hash-table
|
||||
- math
|
||||
- randomized
|
||||
- design
|
||||
status: unsolve
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/insert-delete-getrandom-o1/
|
||||
review_needed: false
|
||||
---
|
||||
@@ -0,0 +1,12 @@
|
||||
---
|
||||
id: 442
|
||||
title: Find All Duplicates in an Array
|
||||
difficulty: Medium
|
||||
tags:
|
||||
- array
|
||||
- hash-table
|
||||
status: unsolve
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/find-all-duplicates-in-an-array/
|
||||
review_needed: false
|
||||
---
|
||||
@@ -0,0 +1,14 @@
|
||||
---
|
||||
id: 49
|
||||
title: Group Anagrams
|
||||
difficulty: Medium
|
||||
tags:
|
||||
- array
|
||||
- hash-table
|
||||
- string
|
||||
- sorting
|
||||
status: unsolve
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/group-anagrams/
|
||||
review_needed: false
|
||||
---
|
||||
@@ -0,0 +1,13 @@
|
||||
---
|
||||
id: 560
|
||||
title: Subarray Sum Equals K
|
||||
difficulty: Medium
|
||||
tags:
|
||||
- array
|
||||
- hash-table
|
||||
- prefix-sum
|
||||
status: unsolve
|
||||
date_solved: 2026-05-26
|
||||
leetcode_url: https://leetcode.com/problems/subarray-sum-equals-k/
|
||||
review_needed: false
|
||||
---
|
||||
@@ -0,0 +1,375 @@
|
||||
---
|
||||
tags:
|
||||
- leetcode/array-hashing
|
||||
- study-guide
|
||||
- coding-practice
|
||||
difficulty_distribution:
|
||||
easy: 7
|
||||
medium: 10
|
||||
hard: 3
|
||||
total_problems: 20
|
||||
completed_problems: 1
|
||||
progress_percentage: 5%
|
||||
last_updated: 2026-05-26
|
||||
---
|
||||
|
||||
# 🚀 Array & Hashing: LeetCode Master List
|
||||
|
||||
Welcome to your study guide for **Array and Hashing**! This topic is the bedrock of coding interviews, establishing the core patterns for lookups, frequency counting, prefix arrays, and space-time tradeoffs.
|
||||
|
||||
---
|
||||
|
||||
## 📊 Progress Tracker
|
||||
|
||||
| Status | # | Problem | Difficulty | Key Technique | Links |
|
||||
| :----: | :--: | :------------------------------------------------------------------------------------------------------------------------ | :--------: | :----------------------------- | :--------------------------------------------------------------------------------------: |
|
||||
| [x] | 1 | [Two Sum](https://leetcode.com/problems/two-sum/) | 🟢 Easy | Hash Map Complement | [LeetCode](https://leetcode.com/problems/two-sum/) |
|
||||
| [ ] | 217 | [Contains Duplicate](https://leetcode.com/problems/contains-duplicate/) | 🟢 Easy | Hash Set Presence | [LeetCode](https://leetcode.com/problems/contains-duplicate/) |
|
||||
| [ ] | 242 | [Valid Anagram](https://leetcode.com/problems/valid-anagram/) | 🟢 Easy | Frequency Count / Sorting | [LeetCode](https://leetcode.com/problems/valid-anagram/) |
|
||||
| [ ] | 169 | [Majority Element](https://leetcode.com/problems/majority-element/) | 🟢 Easy | Boyer-Moore Voting / Hash Map | [LeetCode](https://leetcode.com/problems/majority-element/) |
|
||||
| [ ] | 290 | [Word Pattern](https://leetcode.com/problems/word-pattern/) | 🟢 Easy | Bijective Hash Mapping | [LeetCode](https://leetcode.com/problems/word-pattern/) |
|
||||
| [ ] | 1346 | [Check If N and Its Double Exist](https://leetcode.com/problems/check-if-n-and-its-double-exist/) | 🟢 Easy | Hash Set Double-Lookup | [LeetCode](https://leetcode.com/problems/check-if-n-and-its-double-exist/) |
|
||||
| [ ] | 724 | [Find Pivot Index](https://leetcode.com/problems/find-pivot-index/) | 🟢 Easy | Prefix Sum Balance | [LeetCode](https://leetcode.com/problems/find-pivot-index/) |
|
||||
| [ ] | 49 | [Group Anagrams](https://leetcode.com/problems/group-anagrams/) | 🟡 Medium | Categorization Key Mapping | [LeetCode](https://leetcode.com/problems/group-anagrams/) |
|
||||
| [ ] | 347 | [Top K Frequent Elements](https://leetcode.com/problems/top-k-frequent-elements/) | 🟡 Medium | Bucket Sort / Heap Tracking | [LeetCode](https://leetcode.com/problems/top-k-frequent-elements/) |
|
||||
| [ ] | 238 | [Product of Array Except Self](https://leetcode.com/problems/product-of-array-except-self/) | 🟡 Medium | Left/Right Prefix Products | [LeetCode](https://leetcode.com/problems/product-of-array-except-self/) |
|
||||
| [ ] | 36 | [Valid Sudoku](https://leetcode.com/problems/valid-sudoku/) | 🟡 Medium | Sub-grid Hashing (Bitmask/Set) | [LeetCode](https://leetcode.com/problems/valid-sudoku/) |
|
||||
| [ ] | 128 | [Longest Consecutive Sequence](https://leetcode.com/problems/longest-consecutive-sequence/) | 🟡 Medium | Hash Set Boundary Scan | [LeetCode](https://leetcode.com/problems/longest-consecutive-sequence/) |
|
||||
| [ ] | 560 | [Subarray Sum Equals K](https://leetcode.com/problems/subarray-sum-equals-k/) | 🟡 Medium | Prefix Sum + Frequency Map | [LeetCode](https://leetcode.com/problems/subarray-sum-equals-k/) |
|
||||
| [ ] | 271 | [Encode and Decode Strings](https://leetcode.com/problems/encode-and-decode-strings/) | 🟡 Medium | Length-Prefix Chunking | [LeetCode](https://leetcode.com/problems/encode-and-decode-strings/) |
|
||||
| [ ] | 442 | [Find All Duplicates in an Array](https://leetcode.com/problems/find-all-duplicates-in-an-array/) | 🟡 Medium | In-place Sign Negation | [LeetCode](https://leetcode.com/problems/find-all-duplicates-in-an-array/) |
|
||||
| [ ] | 380 | [Insert Delete GetRandom O(1)](https://leetcode.com/problems/insert-delete-getrandom-o1/) | 🟡 Medium | Array + Map Index Swapping | [LeetCode](https://leetcode.com/problems/insert-delete-getrandom-o1/) |
|
||||
| [ ] | 2017 | [Grid Game](https://leetcode.com/problems/grid-game/) | 🟡 Medium | 2-Row Prefix Sum Selection | [LeetCode](https://leetcode.com/problems/grid-game/) |
|
||||
| [ ] | 41 | [First Missing Positive](https://leetcode.com/problems/first-missing-positive/) | 🔴 Hard | Cyclic In-place Sorting | [LeetCode](https://leetcode.com/problems/first-missing-positive/) |
|
||||
| [ ] | 30 | [Substring with Concatenation](https://leetcode.com/problems/substring-with-concatenation-of-all-words/) | 🔴 Hard | Sliding Window + Count Maps | [LeetCode](https://leetcode.com/problems/substring-with-concatenation-of-all-words/) |
|
||||
| [ ] | 381 | [Insert Delete GetRandom O(1) - Duplicates](https://leetcode.com/problems/insert-delete-getrandom-o1-duplicates-allowed/) | 🔴 Hard | Array + Map to Set of Indices | [LeetCode](https://leetcode.com/problems/insert-delete-getrandom-o1-duplicates-allowed/) |
|
||||
|
||||
---
|
||||
|
||||
## 🟢 Easy Problems (7)
|
||||
|
||||
### 1. Two Sum (LC 1)
|
||||
> [!TIP]
|
||||
> **Core Concept:** Instead of checking all pairs, store the numbers you have seen in a hash map mapping `value -> index`. For each `num`, check if its complement (`target - num`) is already in the map.
|
||||
|
||||
- **Complexity Target:** $O(n)$ Time | $O(n)$ Space
|
||||
- **Starter Template:**
|
||||
```python
|
||||
def twoSum(nums: list[int], target: int) -> list[int]:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
### 2. Contains Duplicate (LC 217)
|
||||
> [!TIP]
|
||||
> **Core Concept:** Iterate through the array and store each element in a Hash Set. If the element is already in the set, a duplicate has been found.
|
||||
|
||||
- **Complexity Target:** $O(n)$ Time | $O(n)$ Space
|
||||
- **Starter Template:**
|
||||
```python
|
||||
def containsDuplicate(nums: list[int]) -> bool:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
### 3. Valid Anagram (LC 242)
|
||||
> [!TIP]
|
||||
> **Core Concept:** Count the frequency of characters in both strings. You can use two hash maps, or a single array size 26 if constraints are purely lowercase a-z. Compare the frequency distributions.
|
||||
|
||||
- **Complexity Target:** $O(n)$ Time | $O(1)$ Space (since lowercase English alphabet count is fixed at 26)
|
||||
- **Starter Template:**
|
||||
```python
|
||||
def isAnagram(s: str, t: str) -> bool:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
### 4. Majority Element (LC 169)
|
||||
> [!TIP]
|
||||
> **Core Concept:** While a Hash Map works, you can solve this in $O(1)$ space using **Boyer-Moore Voting Algorithm**. Keep a `candidate` and a `count`. When `count == 0`, pick the current element as candidate. Increment `count` if the element matches the candidate, else decrement it.
|
||||
|
||||
- **Complexity Target:** $O(n)$ Time | $O(1)$ Space
|
||||
- **Starter Template:**
|
||||
```python
|
||||
def majorityElement(nums: list[int]) -> int:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
### 5. Word Pattern (LC 290)
|
||||
> [!TIP]
|
||||
> **Core Concept:** Establish a bijective (two-way) mapping between character in `pattern` and words in string `s` using two hash maps. If a key maps to a different value in either map, return `False`.
|
||||
|
||||
- **Complexity Target:** $O(n)$ Time | $O(n)$ Space (where $n$ is total chars/words)
|
||||
- **Starter Template:**
|
||||
```python
|
||||
def wordPattern(pattern: str, s: str) -> bool:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
### 6. Check If N and Its Double Exist (LC 1346)
|
||||
> [!TIP]
|
||||
> **Core Concept:** Iterate through the array. Check if `2 * num` or `num / 2` (if divisible by 2) exists in your Hash Set of previously visited values.
|
||||
|
||||
- **Complexity Target:** $O(n)$ Time | $O(n)$ Space
|
||||
- **Starter Template:**
|
||||
```python
|
||||
def checkIfExist(arr: list[int]) -> bool:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
### 7. Find Pivot Index (LC 724)
|
||||
> [!TIP]
|
||||
> **Core Concept:** Compute the total sum of the array. Track the running `left_sum`. For each index, check if `left_sum == total_sum - left_sum - num`. If it is, that's the pivot.
|
||||
|
||||
- **Complexity Target:** $O(n)$ Time | $O(1)$ Space
|
||||
- **Starter Template:**
|
||||
```python
|
||||
def pivotIndex(nums: list[int]) -> int:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
## 🟡 Medium Problems (10)
|
||||
|
||||
### 8. Group Anagrams (LC 49)
|
||||
> [!TIP]
|
||||
> **Core Concept:** Group words by their character signature. The signature can be a sorted string, or a character count tuple `[0] * 26` mapped to list of matching strings.
|
||||
|
||||
- **Complexity Target:** $O(n \cdot m)$ Time (where $m$ is max word length) | $O(n \cdot m)$ Space
|
||||
- **Starter Template:**
|
||||
```python
|
||||
def groupAnagrams(strs: list[str]) -> list[list[str]]:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
### 9. Top K Frequent Elements (LC 347)
|
||||
> [!TIP]
|
||||
> **Core Concept:** Count frequencies using a hash map. Instead of sorting (which is $O(n \log n)$), use **Bucket Sort** where index represents frequencies. Since the max frequency is capped at `len(nums)`, we can assemble the result in linear time.
|
||||
|
||||
- **Complexity Target:** $O(n)$ Time | $O(n)$ Space
|
||||
- **Starter Template:**
|
||||
```python
|
||||
def topKFrequent(nums: list[int], k: int) -> list[int]:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
### 10. Product of Array Except Self (LC 238)
|
||||
> [!TIP]
|
||||
> **Core Concept:** Create an output array. Do a forward pass to store the prefix product at each index. Then do a backward pass, keeping a running suffix product, multiplying it into your result.
|
||||
|
||||
- **Complexity Target:** $O(n)$ Time | $O(1)$ Extra Space (excluding output array)
|
||||
- **Starter Template:**
|
||||
```python
|
||||
def productExceptSelf(nums: list[int]) -> list[int]:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
### 11. Valid Sudoku (LC 36)
|
||||
> [!TIP]
|
||||
> **Core Concept:** Track duplicates in rows, columns, and 3x3 sub-grids. Use a set for each row, column, and sub-grid. The subgrid index can be identified using `(r // 3, c // 3)`.
|
||||
|
||||
- **Complexity Target:** $O(1)$ Time & Space (since grid size is constant $9 \times 9$)
|
||||
- **Starter Template:**
|
||||
```python
|
||||
def isValidSudoku(board: list[list[str]]) -> bool:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
### 12. Longest Consecutive Sequence (LC 128)
|
||||
> [!TIP]
|
||||
> **Core Concept:** Convert the array to a Hash Set. Loop through each number; if `num - 1` is not in the set, it means `num` is the *start* of a new sequence. Scan forward (`num + 1`, `num + 2`...) to find the sequence length. This guarantees each element is processed at most twice.
|
||||
|
||||
- **Complexity Target:** $O(n)$ Time | $O(n)$ Space
|
||||
- **Starter Template:**
|
||||
```python
|
||||
def longestConsecutive(nums: list[int]) -> int:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
### 13. Subarray Sum Equals K (LC 560)
|
||||
> [!TIP]
|
||||
> **Core Concept:** Use prefix sum properties. If the difference between current prefix sum and target `k` (i.e. `prefix_sum - k`) was seen previously as a prefix sum, the subarray between those indices sums to `k`. Store prefix sum frequencies in a map.
|
||||
|
||||
- **Complexity Target:** $O(n)$ Time | $O(n)$ Space
|
||||
- **Starter Template:**
|
||||
```python
|
||||
def subarraySum(nums: list[int], k: int) -> int:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
### 14. Encode and Decode Strings (LC 271 / Premium)
|
||||
> [!TIP]
|
||||
> **Core Concept:** To combine strings safely, prepend each string with its length followed by a delimiter (e.g. `"4#neet"`). When decoding, parse the number, jump past the delimiter, and slice that exact length.
|
||||
|
||||
- **Complexity Target:** $O(n)$ Time | $O(1)$ Extra Space (excluding output list)
|
||||
- **Starter Template:**
|
||||
```python
|
||||
class Codec:
|
||||
def encode(self, strs: list[str]) -> str:
|
||||
pass
|
||||
|
||||
def decode(self, s: str) -> list[str]:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
### 15. Find All Duplicates in an Array (LC 442)
|
||||
> [!TIP]
|
||||
> **Core Concept:** The array contains elements from `1` to `n`. You can use the values as indices. As you iterate, look up `nums[abs(x) - 1]`. If it is positive, negate it. If it is already negative, then `abs(x)` has been seen before.
|
||||
|
||||
- **Complexity Target:** $O(n)$ Time | $O(1)$ Extra Space
|
||||
- **Starter Template:**
|
||||
```python
|
||||
def findDuplicates(nums: list[int]) -> list[int]:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
### 16. Insert Delete GetRandom O(1) (LC 380)
|
||||
> [!TIP]
|
||||
> **Core Concept:** Store values in an Array List to achieve $O(1)$ random lookups. Maintain a Hash Map `val -> index` to achieve $O(1)$ insertions and updates. When deleting, swap the element to delete with the last element in the array, update the map, and pop from the array.
|
||||
|
||||
- **Complexity Target:** $O(1)$ Time average | $O(n)$ Space
|
||||
- **Starter Template:**
|
||||
```python
|
||||
class RandomizedSet:
|
||||
def __init__(self):
|
||||
pass
|
||||
|
||||
def insert(self, val: int) -> bool:
|
||||
pass
|
||||
|
||||
def remove(self, val: int) -> bool:
|
||||
pass
|
||||
|
||||
def getRandom(self) -> int:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
### 17. Grid Game (LC 2017)
|
||||
> [!TIP]
|
||||
> **Core Concept:** The first robot splits the grid into two paths. Compute prefix sums for row 0 and row 1. Robot 1 wants to minimize Robot 2's maximum score, which will always be the max of the remaining top-right elements or bottom-left elements after Robot 1 pivots.
|
||||
|
||||
- **Complexity Target:** $O(n)$ Time | $O(n)$ Space (or $O(1)$ if prefix sums are computed on-the-fly)
|
||||
- **Starter Template:**
|
||||
```python
|
||||
def gridGame(grid: list[list[int]]) -> int:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
## 🔴 Hard Problems (3)
|
||||
|
||||
### 18. First Missing Positive (LC 41)
|
||||
> [!TIP]
|
||||
> **Core Concept:** Use **Cyclic Sort** style placement. Iterate through the array and try to place each number `x` (if it lies in range `[1, n]`) at its correct index `x - 1`. Perform this swap in a loop. Finally, scan the array to find the first index `i` where `nums[i] != i + 1`.
|
||||
|
||||
- **Complexity Target:** $O(n)$ Time | $O(1)$ Space
|
||||
- **Starter Template:**
|
||||
```python
|
||||
def firstMissingPositive(nums: list[int]) -> int:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
### 19. Substring with Concatenation of All Words (LC 30)
|
||||
> [!TIP]
|
||||
> **Core Concept:** All words are of equal length `L`. Use a sliding window starting at each offset `0 <= i < L`. For each window, use a hash map to count occurrences of words of length `L` and compare it with the count map of the target `words` list.
|
||||
|
||||
- **Complexity Target:** $O(n \cdot L)$ Time (where $n$ is string length, $L$ is word length) | $O(m \cdot L)$ Space (where $m$ is number of words)
|
||||
- **Starter Template:**
|
||||
```python
|
||||
def findSubstring(s: str, words: list[str]) -> list[int]:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
|
||||
---
|
||||
|
||||
### 20. Insert Delete GetRandom O(1) - Duplicates Allowed (LC 381)
|
||||
> [!TIP]
|
||||
> **Core Concept:** Similar to LC 380, but the Hash Map now maps `val -> Set of indices` where the value resides in our dynamic list. During deletion, lookup any index from the value's set, swap with the list's last element, and update the set of the swapped element accordingly.
|
||||
|
||||
- **Complexity Target:** $O(1)$ Time average | $O(n)$ Space
|
||||
- **Starter Template:**
|
||||
```python
|
||||
class RandomizedCollection:
|
||||
def __init__(self):
|
||||
pass
|
||||
|
||||
def insert(self, val: int) -> bool:
|
||||
pass
|
||||
|
||||
def remove(self, val: int) -> bool:
|
||||
pass
|
||||
|
||||
def getRandom(self) -> int:
|
||||
pass
|
||||
```
|
||||
- **My Notes / Solution:**
|
||||
*
|
||||
@@ -0,0 +1,237 @@
|
||||
---
|
||||
title: Introduction to NumPy
|
||||
tags:
|
||||
- python
|
||||
- numpy
|
||||
- data-science
|
||||
- numerical-computing
|
||||
category: Lesson
|
||||
difficulty: Beginner to Intermediate
|
||||
created: 2026-05-26
|
||||
---
|
||||
|
||||
# Introduction to NumPy
|
||||
|
||||
> [!abstract] What is NumPy?
|
||||
> **NumPy** (Numerical Python) is the foundational library for scientific computing in Python. It provides a high-performance multidimensional array object, tools for working with these arrays, and linear algebra, Fourier transform, and random number capabilities.
|
||||
>
|
||||
> ### Why use NumPy instead of standard Python lists?
|
||||
> 1. **Speed:** NumPy arrays are written in C, making mathematical operations up to 100x faster than standard Python lists.
|
||||
> 2. **Memory Efficiency:** NumPy arrays use contiguous blocks of memory, whereas Python lists store pointers to objects scattered across memory.
|
||||
> 3. **Vectorization:** It allows performing mathematical operations on whole arrays without writing slow `for` loops.
|
||||
|
||||
---
|
||||
|
||||
## 📐 The N-Dimensional Array (`ndarray`)
|
||||
|
||||
The core of NumPy is the **`ndarray`** (N-dimensional array). It is a grid of values, all of the **same type** (homogenous), indexed by a tuple of non-negative integers.
|
||||
|
||||
```mermaid
|
||||
graph TD
|
||||
A[ndarray] --> B["1D Array (Vector) <br> Shape: (n,)"]
|
||||
A --> C["2D Array (Matrix) <br> Shape: (m, n)"]
|
||||
A --> D["3D Array (Tensor) <br> Shape: (p, m, n)"]
|
||||
```
|
||||
|
||||
### Essential Array Attributes
|
||||
Every array has attributes that describe its structure:
|
||||
|
||||
```python
|
||||
import numpy as np
|
||||
|
||||
arr = np.array([[1, 2, 3], [4, 5, 6]])
|
||||
|
||||
print(arr.ndim) # Number of dimensions (axes) -> 2
|
||||
print(arr.shape) # Tuple representing sizes in each dimension -> (2, 3)
|
||||
print(arr.size) # Total number of elements -> 6
|
||||
print(arr.dtype) # Data type of the elements -> int64
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🛠️ Creating Arrays
|
||||
|
||||
First, import the library using the standard alias:
|
||||
```python
|
||||
import numpy as np
|
||||
```
|
||||
|
||||
### 1. From Python Lists
|
||||
```python
|
||||
# 1D Vector
|
||||
v = np.array([1, 2, 3])
|
||||
|
||||
# 2D Matrix
|
||||
m = np.array([[1, 2], [3, 4]])
|
||||
```
|
||||
|
||||
### 2. Built-in Placeholders
|
||||
NumPy provides functions to initialize arrays with placeholders, avoiding manual creation:
|
||||
|
||||
```python
|
||||
# Array of zeros
|
||||
zeros = np.zeros((3, 4)) # 3 rows, 4 columns
|
||||
|
||||
# Array of ones
|
||||
ones = np.ones((2, 3), dtype=np.int32)
|
||||
|
||||
# Range of numbers (similar to range())
|
||||
range_arr = np.arange(0, 10, 2) # [0, 2, 4, 6, 8]
|
||||
|
||||
# Linearly spaced numbers
|
||||
linspace_arr = np.linspace(0, 1, 5) # [0.0, 0.25, 0.5, 0.75, 1.0]
|
||||
|
||||
# Identity Matrix
|
||||
eye_matrix = np.eye(3) # 3x3 identity matrix
|
||||
```
|
||||
|
||||
### 3. Random Number Generation
|
||||
```python
|
||||
# Uniform random values between [0.0, 1.0)
|
||||
rand_arr = np.random.rand(2, 2)
|
||||
|
||||
# Standard normal distribution (mean=0, std=1)
|
||||
randn_arr = np.random.randn(2, 2)
|
||||
|
||||
# Random integers
|
||||
rand_ints = np.random.randint(1, 100, size=(5,))
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## ⚡ Element-wise Operations & Vectorization
|
||||
|
||||
In standard Python, to add two lists element-wise, you need a list comprehension or loop. In NumPy, you do it directly.
|
||||
|
||||
```python
|
||||
x = np.array([1, 2, 3])
|
||||
y = np.array([4, 5, 6])
|
||||
|
||||
print(x + y) # [5, 7, 9]
|
||||
print(x * y) # [4, 10, 18]
|
||||
print(x ** 2) # [1, 4, 9]
|
||||
```
|
||||
|
||||
### 📡 Broadcasting
|
||||
Broadcasting is a powerful mechanism that allows NumPy to perform arithmetic operations on arrays of **different shapes**. The smaller array is "broadcast" across the larger array so that they have compatible shapes.
|
||||
|
||||
```python
|
||||
matrix = np.array([[1, 2, 3], [4, 5, 6]])
|
||||
scalar = 10
|
||||
|
||||
# The scalar is added to every single element
|
||||
print(matrix + scalar)
|
||||
# [[11, 12, 13]
|
||||
# [14, 15, 16]]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🔍 Indexing, Slicing & Masking
|
||||
|
||||
### Slicing 2D Arrays
|
||||
Slicing follows the format `array[row_start:row_end, col_start:col_end]`.
|
||||
|
||||
```python
|
||||
arr = np.array([
|
||||
[10, 11, 12],
|
||||
[20, 21, 22],
|
||||
[30, 31, 32]
|
||||
])
|
||||
|
||||
# Get row at index 1
|
||||
print(arr[1, :]) # [20, 21, 22]
|
||||
|
||||
# Get column at index 2
|
||||
print(arr[:, 2]) # [12, 22, 32]
|
||||
|
||||
# Slice a subgrid (top-left 2x2)
|
||||
print(arr[0:2, 0:2])
|
||||
# [[10, 11]
|
||||
# [20, 21]]
|
||||
```
|
||||
|
||||
### 🎭 Boolean Masking (Conditional Filtering)
|
||||
You can filter arrays using conditions. NumPy returns elements where the condition resolves to `True`.
|
||||
|
||||
```python
|
||||
data = np.array([1, 5, 8, 12, 3, 15])
|
||||
|
||||
# Create a boolean mask
|
||||
mask = data > 5 # [False, False, True, True, False, True]
|
||||
|
||||
# Filter using the mask
|
||||
filtered_data = data[mask] # [8, 12, 15]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🧮 Common Aggregations & Axis Operations
|
||||
|
||||
Aggregations allow you to compute statistics over entire arrays or along specific **axes**:
|
||||
* `axis=0`: Down the columns (collapses rows).
|
||||
* `axis=1`: Across the rows (collapses columns).
|
||||
|
||||
```python
|
||||
arr = np.array([[1, 2], [3, 4]])
|
||||
|
||||
# Sum of all elements
|
||||
print(np.sum(arr)) # 10
|
||||
|
||||
# Sum down the columns (vertical)
|
||||
print(np.sum(arr, axis=0)) # [4, 6]
|
||||
|
||||
# Sum across the rows (horizontal)
|
||||
print(np.sum(arr, axis=1)) # [3, 7]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🔄 Reshaping and Transposing
|
||||
|
||||
You can change the shape of an array without changing its data using `.reshape()` or `.T` (Transpose).
|
||||
|
||||
```python
|
||||
flat = np.arange(1, 7) # [1, 2, 3, 4, 5, 6]
|
||||
|
||||
# Reshape into a 2x3 matrix
|
||||
matrix = flat.reshape(2, 3)
|
||||
# [[1, 2, 3]
|
||||
# [4, 5, 6]]
|
||||
|
||||
# Transpose matrix (swap rows and columns)
|
||||
transposed = matrix.T
|
||||
# [[1, 4]
|
||||
# [2, 5]
|
||||
# [3, 6]]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 💡 Best Practices
|
||||
|
||||
> [!important] Avoid Standard Loops
|
||||
> Standard loops in Python are interpreted, which adds massive overhead. Vectorized operations execute in compiled C, taking advantage of CPU caches and SIMD instructions.
|
||||
>
|
||||
> **Example Comparison:**
|
||||
> ```python
|
||||
> # ❌ Extremely Slow
|
||||
> values = np.random.rand(1_000_000)
|
||||
> reciprocal = [1 / x for x in values]
|
||||
>
|
||||
> # ✅ Near Instantaneous
|
||||
> reciprocal = 1 / values
|
||||
> ```
|
||||
|
||||
> [!tip] Use In-place Operations to Save Memory
|
||||
> Instead of creating a new copy, perform calculations directly on the existing array if possible using syntax like `+=`, `-=`, or `*=`.
|
||||
> ```python
|
||||
> a = np.ones(1000000)
|
||||
> b = np.ones(1000000)
|
||||
>
|
||||
> # Allocates new memory
|
||||
> a = a + b
|
||||
>
|
||||
> # Modifies 'a' in-place (saves memory allocation time)
|
||||
> a += b
|
||||
> ```
|
||||
@@ -0,0 +1,241 @@
|
||||
---
|
||||
title: Introduction to Pandas
|
||||
tags:
|
||||
- python
|
||||
- pandas
|
||||
- data-science
|
||||
- data-analysis
|
||||
category: Lesson
|
||||
difficulty: Beginner to Intermediate
|
||||
created: 2026-05-26
|
||||
---
|
||||
|
||||
# Introduction to Pandas
|
||||
|
||||
> [!abstract] What is Pandas?
|
||||
> **Pandas** is a fast, powerful, flexible, and easy-to-use open-source data analysis and manipulation library built on top of the Python programming language.
|
||||
>
|
||||
> The name is derived from **"Panel Data"**, an econometrics term for multidimensional structured data sets. It is the foundational library for Data Science, Machine Learning, and Data Analysis in Python, acting as the bridge between raw data files (like CSVs, Excel files, or SQL databases) and numerical/modeling libraries (like NumPy, Scikit-Learn, and PyTorch).
|
||||
|
||||
---
|
||||
|
||||
## 🏗️ Core Data Structures
|
||||
|
||||
Pandas simplifies data manipulation by providing two primary, highly optimized data structures:
|
||||
|
||||
```mermaid
|
||||
graph TD
|
||||
A[Pandas Data Structures] --> B[Series - 1D]
|
||||
A --> C[DataFrame - 2D]
|
||||
B -->|Multiple columns merged| C
|
||||
C -->|Single column extracted| B
|
||||
```
|
||||
|
||||
### 1. Series (1D)
|
||||
A **Series** is a one-dimensional array-like object containing an array of data and an associated array of data labels, called its **index**. Think of it as a single column in a spreadsheet.
|
||||
|
||||
```python
|
||||
import pandas as pd
|
||||
|
||||
# Creating a Series
|
||||
temperatures = pd.Series([22.5, 24.0, 19.5, 21.8], name="Temp")
|
||||
print(temperatures)
|
||||
```
|
||||
|
||||
### 2. DataFrame (2D)
|
||||
A **DataFrame** represents a tabular, spreadsheet-like data structure containing an ordered collection of columns, each of which can be a different value type (numeric, string, boolean, etc.). It has both a row index and a column index.
|
||||
|
||||
| Index | Name | Age | Department |
|
||||
| :--- | :--- | :--- | :--- |
|
||||
| **0** | Alice | 28 | Engineering |
|
||||
| **1** | Bob | 34 | Marketing |
|
||||
| **2** | Charlie | 22 | HR |
|
||||
|
||||
---
|
||||
|
||||
## 🛠️ Getting Started & Creating Data
|
||||
|
||||
First, make sure you have pandas imported. The universal convention is to alias it as `pd`.
|
||||
|
||||
```python
|
||||
import pandas as pd
|
||||
import numpy as np
|
||||
```
|
||||
|
||||
### Creating DataFrames Manually
|
||||
You can easily create DataFrames from dictionaries or lists:
|
||||
|
||||
```python
|
||||
data = {
|
||||
'Name': ['Alice', 'Bob', 'Charlie', 'David'],
|
||||
'Age': [25, 30, 35, 40],
|
||||
'Salary': [70000, 80000, 120000, 90000],
|
||||
'Department': ['HR', 'Engineering', 'Engineering', 'Finance']
|
||||
}
|
||||
|
||||
df = pd.DataFrame(data)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🔍 Essential Operations (The "Big Five")
|
||||
|
||||
### 1. Reading & Writing Data
|
||||
Pandas supports a wide range of file formats out of the box.
|
||||
|
||||
```python
|
||||
# Reading data
|
||||
df_csv = pd.read_csv('employees.csv')
|
||||
df_excel = pd.read_excel('sales.xlsx', sheet_name='Sheet1')
|
||||
df_sql = pd.read_sql('SELECT * FROM users', database_connection)
|
||||
|
||||
# Writing data
|
||||
df.to_csv('output.csv', index=False) # index=False avoids writing the row numbers
|
||||
df.to_json('output.json')
|
||||
```
|
||||
|
||||
### 2. Inspecting Your Data
|
||||
Before doing any analysis, you must understand the shape and types of your data.
|
||||
|
||||
```python
|
||||
df.head(2) # Returns the first 2 rows
|
||||
df.tail(2) # Returns the last 2 rows
|
||||
df.info() # Summary of columns, non-null counts, and data types
|
||||
df.describe() # Generates descriptive statistics for numerical columns
|
||||
df.shape # Returns (rows, columns) as a tuple
|
||||
```
|
||||
|
||||
### 3. Selection & Filtering
|
||||
Selecting data is one of the most common tasks. Pandas provides multiple intuitive ways to do this.
|
||||
|
||||
#### Column Selection
|
||||
```python
|
||||
# Select a single column (returns a Series)
|
||||
ages = df['Age']
|
||||
|
||||
# Select multiple columns (returns a DataFrame)
|
||||
subset = df[['Name', 'Salary']]
|
||||
```
|
||||
|
||||
#### Row Selection using `.loc` and `.iloc`
|
||||
* `.loc` is **label-based**: references rows/columns by their row labels or column names.
|
||||
* `.iloc` is **integer-position-based**: references rows/columns by their 0-indexed positions.
|
||||
|
||||
```python
|
||||
# Select the first row by position
|
||||
first_row = df.iloc[0]
|
||||
|
||||
# Select a cell by row label and column name
|
||||
val = df.loc[2, 'Name'] # Charlie
|
||||
```
|
||||
|
||||
#### Boolean Indexing (Filtering)
|
||||
To filter rows based on conditions:
|
||||
|
||||
```python
|
||||
# Filter employees earning more than 85,000
|
||||
high_earners = df[df['Salary'] > 85000]
|
||||
|
||||
# Combine multiple conditions using & (AND) or | (OR)
|
||||
# Always wrap conditions in parentheses!
|
||||
eng_seniors = df[(df['Department'] == 'Engineering') & (df['Age'] > 30)]
|
||||
```
|
||||
|
||||
### 4. Data Cleaning
|
||||
Raw data is rarely perfect. Pandas excels at handling missing values and data type conversions.
|
||||
|
||||
```python
|
||||
# Check for missing values
|
||||
df.isna().sum()
|
||||
|
||||
# Drop rows with missing values
|
||||
df_clean = df.dropna()
|
||||
|
||||
# Fill missing values with a default/placeholder
|
||||
df['Salary'] = df['Salary'].fillna(df['Salary'].mean())
|
||||
|
||||
# Rename columns
|
||||
df = df.rename(columns={'Name': 'Full Name', 'Age': 'Years'})
|
||||
```
|
||||
|
||||
### 5. Grouping & Aggregation
|
||||
To summarize data by categories, use the Split-Apply-Combine workflow via `groupby`.
|
||||
|
||||
```python
|
||||
# Calculate the average salary by department
|
||||
dept_salaries = df.groupby('Department')['Salary'].mean()
|
||||
|
||||
# Perform multiple aggregations at once
|
||||
summary = df.groupby('Department').agg({
|
||||
'Salary': ['mean', 'min', 'max'],
|
||||
'Age': 'mean'
|
||||
})
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🎓 Hands-on Practice Walkthrough
|
||||
|
||||
Let's walk through a mini-scenario. Imagine we have the following DataFrame of store transactions:
|
||||
|
||||
```python
|
||||
transactions = pd.DataFrame({
|
||||
'TransactionID': [101, 102, 103, 104, 105],
|
||||
'Store': ['North', 'South', 'North', 'West', 'South'],
|
||||
'Amount': [250.50, 150.00, np.nan, 300.25, 450.00],
|
||||
'ItemCount': [3, 2, 1, 5, 4]
|
||||
})
|
||||
```
|
||||
|
||||
Let's answer three questions:
|
||||
|
||||
### Question 1: Fill missing amounts with the median transaction amount.
|
||||
```python
|
||||
median_amount = transactions['Amount'].median() # 275.375
|
||||
transactions['Amount'] = transactions['Amount'].fillna(median_amount)
|
||||
```
|
||||
|
||||
### Question 2: Find transactions with a total amount greater than 200.
|
||||
```python
|
||||
large_tx = transactions[transactions['Amount'] > 200]
|
||||
```
|
||||
|
||||
### Question 3: Find the total revenue (sum of amounts) generated by each store.
|
||||
```python
|
||||
store_revenue = transactions.groupby('Store')['Amount'].sum().reset_index()
|
||||
print(store_revenue)
|
||||
# Store Amount
|
||||
# 0 North 525.875
|
||||
# 1 South 600.00
|
||||
# 2 West 300.25
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## ⚡ Performance Tips & Best Practices
|
||||
|
||||
> [!warning] Don't Loop Over Rows!
|
||||
> Never write `for index, row in df.iterrows():` unless absolutely necessary. Iteration is extremely slow because it disables pandas' vectorized backend.
|
||||
>
|
||||
> **Instead, use Vectorization:**
|
||||
> ```python
|
||||
> # ❌ Slow & non-pythonic
|
||||
> for i in range(len(df)):
|
||||
> df.loc[i, 'Tax'] = df.loc[i, 'Salary'] * 0.1
|
||||
>
|
||||
> # ✅ Fast & vectorised
|
||||
> df['Tax'] = df['Salary'] * 0.1
|
||||
> ```
|
||||
|
||||
> [!tip] Avoid the SettingWithCopyWarning
|
||||
> When you slice a DataFrame and then modify it, pandas warns you that you might be editing a temporary copy rather than the original source.
|
||||
> To prevent this, use `.copy()` when creating a subset you intend to modify:
|
||||
> ```python
|
||||
> # ❌ Might trigger SettingWithCopyWarning
|
||||
> engineering = df[df['Department'] == 'Engineering']
|
||||
> engineering['Bonus'] = 1000
|
||||
>
|
||||
> # ✅ Clean and safe
|
||||
> engineering = df[df['Department'] == 'Engineering'].copy()
|
||||
> engineering['Bonus'] = 1000
|
||||
> ```
|
||||
@@ -0,0 +1,130 @@
|
||||
---
|
||||
title: Learn Machine Learning Like a GENIUS and Not Waste Time
|
||||
channel: InfiniteCodes
|
||||
url: https://youtu.be/i_LwzRVP7bg
|
||||
publish_date: 2024-11-14
|
||||
tags:
|
||||
- machine-learning
|
||||
- study-methodology
|
||||
- learning-strategies
|
||||
- productivity
|
||||
category: Video Summary
|
||||
rating: 5/5
|
||||
---
|
||||
|
||||
# Learn Machine Learning Like a GENIUS and Not Waste Time
|
||||
|
||||
> [!abstract] Executive Summary
|
||||
> This note summarizes the video guide **"Learn Machine Learning Like a GENIUS and Not Waste Time"** by *InfiniteCodes*. The core message is that mastering machine learning (ML) isn't about memorizing complex models or chasing every new trend. Instead, it relies on building a deep intuition of the fundamentals, practicing active learning through project building, reading official documentation, and avoiding the traps of "tutorial hell" and "vibe coding" (blindly relying on LLMs).
|
||||
|
||||
---
|
||||
|
||||
## 🗺️ The ML Learning Roadmap (Order of Operations)
|
||||
|
||||
Beginners often rush into complex deep learning models before understanding basic concepts. The video advocates for a strict, logical **Order of Operations**:
|
||||
|
||||
```mermaid
|
||||
graph TD
|
||||
A[1. Mathematics Foundations] -->|Linear Algebra & Calculus| B[2. Exploratory Data Analysis]
|
||||
B -->|Clean, Wrangling, Feature Eng| C[3. Simple Models First]
|
||||
C -->|Linear/Logistic Reg, Trees| D[4. Advanced Architectures]
|
||||
D -->|CNNs, RNNs, Transformers| E[5. Real-world Deployment]
|
||||
```
|
||||
|
||||
### 1. Mathematics Foundations
|
||||
You don't need a math PhD, but you must build a working intuition of:
|
||||
* **Linear Algebra**: Understanding how data is represented and manipulated as matrices and vectors.
|
||||
* **Calculus**: Grasping **Gradient Descent** (derivatives, partial derivatives) as the optimization engine that allows models to learn.
|
||||
|
||||
### 2. Exploratory Data Analysis (EDA)
|
||||
Before fitting any model, you must "interview" your data:
|
||||
* Use libraries like `pandas` and `numpy` to clean and structure data.
|
||||
* Perform feature engineering and visualize distributions to identify patterns.
|
||||
|
||||
### 3. Simple Models First
|
||||
Start with highly interpretable, foundational algorithms:
|
||||
* **Linear Regression** & **Logistic Regression**
|
||||
* **Decision Trees**
|
||||
* *Why?* They are faster to train, easier to debug, and provide a baseline for more complex models.
|
||||
|
||||
### 4. Advanced Architectures
|
||||
Only dive into deep learning (CNNs, RNNs, Transformers, etc.) when your specific project requires it and your foundations are solid.
|
||||
|
||||
---
|
||||
|
||||
## 🧠 The "Genius" Implementation Strategy
|
||||
|
||||
To truly master an algorithm, do not just read about it. Use this three-step implementation loop:
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Step1[1. Code from Scratch] --> Step2[2. Use Library]
|
||||
Step2 --> Step3[3. Apply to Real Data]
|
||||
```
|
||||
|
||||
1. **Code from Scratch**: Implement the core algorithm in raw Python (using only `numpy`) to understand the underlying mathematics.
|
||||
2. **Use a Library**: Implement the same algorithm using `scikit-learn` to see how it is optimized and structured in production-ready libraries.
|
||||
3. **Apply to Real Data**: Train both implementations on a dataset you gathered or prepared yourself (avoiding clean "toy" datasets).
|
||||
|
||||
> [!example] From Scratch vs. Library Example (Linear Regression)
|
||||
> Below is a comparison of how you should study an algorithm.
|
||||
>
|
||||
> === "From Scratch (Math Intuition)"
|
||||
> ```python
|
||||
> import numpy as np
|
||||
>
|
||||
> class SimpleLinearRegression:
|
||||
> def __init__(self, lr=0.01, epochs=1000):
|
||||
> self.lr = lr
|
||||
> self.epochs = epochs
|
||||
> self.weights = None
|
||||
> self.bias = None
|
||||
>
|
||||
> def fit(self, X, y):
|
||||
> n_samples, n_features = X.shape
|
||||
> self.weights = np.zeros(n_features)
|
||||
> self.bias = 0
|
||||
>
|
||||
> # Gradient Descent loop
|
||||
> for _ in range(self.epochs):
|
||||
> y_predicted = np.dot(X, self.weights) + self.bias
|
||||
> dw = (1 / n_samples) * np.dot(X.T, (y_predicted - y))
|
||||
> db = (1 / n_samples) * np.sum(y_predicted - y)
|
||||
>
|
||||
> self.weights -= self.lr * dw
|
||||
> self.bias -= self.lr * db
|
||||
> ```
|
||||
>
|
||||
> === "Using a Library (Production Standard)"
|
||||
> ```python
|
||||
> from sklearn.linear_model import LinearRegression
|
||||
>
|
||||
> # Initialize and fit
|
||||
> model = LinearRegression()
|
||||
> model.fit(X_train, y_train)
|
||||
>
|
||||
> # Make predictions
|
||||
> predictions = model.predict(X_test)
|
||||
> ```
|
||||
|
||||
---
|
||||
|
||||
## 🚫 Key Pitfalls to Avoid
|
||||
|
||||
> [!danger] 1. The "3-Month Fallacy"
|
||||
> Avoid courses or tutorials promising machine learning mastery in 3 months. Transitioning to a professional level takes sustained, long-term effort and continuous learning.
|
||||
|
||||
> [!warning] 2. Tutorial Hell & "Vibe Coding"
|
||||
> * **Tutorial Hell**: Mindlessly consuming tutorials without writing code. Rule of thumb: Watch **maximum 2 tutorials** on a topic, then immediately build something yourself.
|
||||
> * **Vibe Coding**: Relying entirely on LLMs (like Cursor or Copilot) to generate code for you. If you don't understand the lines of code being generated, you are building a house of cards.
|
||||
|
||||
> [!important] 3. Documentation First
|
||||
> Build a habit of reading the **official documentation** (e.g., PyTorch, scikit-learn, Pandas) instead of asking an AI for code snippets immediately. This builds strong neural connections and developer independence.
|
||||
|
||||
---
|
||||
|
||||
## ⚡ Soft Skills & Mindset
|
||||
|
||||
* **Deep Work**: Dedicate uninterrupted **90 to 120-minute** blocks to study and code. Turn off notifications and focus deeply.
|
||||
* **Problem Solving**: ML is about breaking complex, abstract real-world problems down into structured data steps.
|
||||
* **Community & Networking**: Share your learning journey on GitHub, LinkedIn, or community Discords. Learning in public accelerates growth and opens career opportunities.
|
||||
@@ -0,0 +1,20 @@
|
||||
---
|
||||
title: LLM Notes Index
|
||||
created: 2026-05-26
|
||||
tags:
|
||||
- llm
|
||||
- index
|
||||
- moc
|
||||
category: index
|
||||
---
|
||||
# LLM Notes Index
|
||||
|
||||
All notes under `note/LLM/`, grouped by subfolder.
|
||||
|
||||
```dataview
|
||||
LIST rows.file.link
|
||||
FROM "note/LLM"
|
||||
WHERE file.name != "index"
|
||||
GROUP BY file.folder
|
||||
SORT file.folder ASC
|
||||
```
|
||||
@@ -0,0 +1,48 @@
|
||||
---
|
||||
title: Learn Linux - The Full Course
|
||||
channel: Boot dev
|
||||
url: https://youtu.be/v392lEyM29A?si=2QnHdGS-DdjtS5Bg
|
||||
tags:
|
||||
- linux
|
||||
- course
|
||||
- devops
|
||||
- terminal
|
||||
date: 2026-05-29
|
||||
---
|
||||
|
||||
# Learn Linux - The Full Course
|
||||
|
||||
## Overview
|
||||
This comprehensive course, led by Lane from **Boot.dev**, is designed to provide developers, DevOps engineers, and IT professionals with a solid foundation in Linux and Unix-like systems. The course focuses on moving beyond basic usage to mastering the command line, navigating filesystems, and using powerful CLI tools.
|
||||
|
||||
## Key Topics
|
||||
|
||||
### Terminals and Shells (04:03)
|
||||
- Deep dive into the fundamental interface of Linux.
|
||||
- Explaining the differences between terminals and shells.
|
||||
- How they facilitate system interaction.
|
||||
|
||||
### Filesystems (20:43)
|
||||
- Navigating the Linux directory structure.
|
||||
- Managing files and understanding data organization on disk.
|
||||
|
||||
### Permissions (51:18)
|
||||
- Critical look at the Linux security model.
|
||||
- User and group management.
|
||||
- Reading and modifying file permissions.
|
||||
|
||||
### Programs (01:13:31)
|
||||
- How software is executed in a Linux environment.
|
||||
- Process management.
|
||||
- Configuration of the system `PATH`.
|
||||
|
||||
### Input/Output (01:38:33)
|
||||
- Mastery of data manipulation using standard input/output.
|
||||
- Redirection and pipes.
|
||||
- Utilities like `grep` and `find`.
|
||||
|
||||
### Packages (02:18:43)
|
||||
- Installing, updating, and managing software dependencies and packages.
|
||||
|
||||
## Conclusion
|
||||
The course transforms the command line from a source of intimidation into a powerful tool for productivity and system control.
|
||||
@@ -0,0 +1,111 @@
|
||||
---
|
||||
tags: [neovim, lazyvim, coding-tools, ide, tutorial]
|
||||
created: 2026-05-29
|
||||
status: complete
|
||||
type: lesson
|
||||
---
|
||||
|
||||
# Master Class: In-Depth Guide to LazyVim
|
||||
|
||||
LazyVim is not just a configuration; it is a **Neovim setup framework** designed to provide a high-performance, IDE-like experience while remaining modular and easy to customize. It leverages the power of `lazy.nvim` to ensure that your editor starts instantly by loading components only when they are needed.
|
||||
|
||||
---
|
||||
|
||||
## 1. The Core Philosophy
|
||||
LazyVim is built on three pillars:
|
||||
1. **Speed:** Everything is lazy-loaded. If you aren't editing a Python file, the Python LSP doesn't load.
|
||||
2. **Sane Defaults:** It comes pre-configured with industry-standard settings for UI, indentation, and search.
|
||||
3. **Modularity:** It separates its core logic from your user configuration, allowing you to update the framework without breaking your personal tweaks.
|
||||
|
||||
---
|
||||
|
||||
## 2. The Integrated Toolbox
|
||||
LazyVim integrates several powerful tools into a cohesive workflow:
|
||||
|
||||
### A. **lazy.nvim (The Engine)**
|
||||
The heart of the system. It manages plugin installation, updates, and lazy-loading.
|
||||
- **Command:** `:Lazy`
|
||||
- **Capabilities:** Check for updates, profile startup time, and manage plugin states.
|
||||
|
||||
### B. **Mason.nvim (The Tool Manager)**
|
||||
A "package manager" for your external dependencies.
|
||||
- **Command:** `:Mason`
|
||||
- **Capabilities:** Easily install and manage LSP servers, DAP (debuggers), linters, and formatters directly from within Neovim.
|
||||
|
||||
### C. **nvim-treesitter (The Parser)**
|
||||
Provides high-performance syntax highlighting and code understanding.
|
||||
- **Command:** `:TSUpdate`
|
||||
- **Capabilities:** Better highlighting, indentation, and "incremental selection" (selecting code blocks logically).
|
||||
|
||||
### D. **Telescope.nvim / fzf-lua (The Searcher)**
|
||||
A fuzzy finder that allows you to find anything in your project.
|
||||
- **Keybinds:** `<leader>ff` (files), `<leader>/` (live grep), `<leader>sk` (keymaps).
|
||||
|
||||
---
|
||||
|
||||
## 3. Essential Keybindings & Workflow
|
||||
LazyVim uses the `<Space>` key as the **Leader**. One of its best features is `which-key.nvim`, which displays a popup showing available commands whenever you press your leader key.
|
||||
|
||||
### **Navigation**
|
||||
- `<leader>e`: Toggle **Neo-tree** (File Explorer).
|
||||
- `H` / `L`: Quickly cycle through open buffers (tabs).
|
||||
- `<leader>bb`: Switch between open buffers.
|
||||
- `<leader>fT`: Open a floating terminal.
|
||||
|
||||
### **Coding & LSP**
|
||||
- `K`: Hover documentation (show what a function/variable does).
|
||||
- `gd`: Go to definition.
|
||||
- `gr`: Go to references.
|
||||
- `<leader>ca`: **Code Actions** (Fixes, imports, refactors).
|
||||
- `<leader>cr`: Rename the symbol under the cursor project-wide.
|
||||
- `[d` / `]d`: Jump to previous/next error or warning.
|
||||
|
||||
### **Git Integration**
|
||||
- `<leader>gg`: Open **LazyGit** (a full TUI for Git).
|
||||
- `<leader>gj`: Next Git hunk.
|
||||
- `<leader>gk`: Previous Git hunk.
|
||||
|
||||
---
|
||||
|
||||
## 4. Customizing Your Setup
|
||||
LazyVim's structure is designed to be clean:
|
||||
- `lua/config/options.lua`: Global Neovim settings (e.g., line numbers, tab widths).
|
||||
- `lua/config/keymaps.lua`: Your custom keyboard shortcuts.
|
||||
- `lua/plugins/`: Any `.lua` file created here is automatically loaded as a plugin configuration.
|
||||
|
||||
### **Adding a Plugin**
|
||||
Create `lua/plugins/example.lua`:
|
||||
```lua
|
||||
return {
|
||||
"username/plugin-name",
|
||||
opts = {
|
||||
-- plugin configuration goes here
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
### **Enabling "Extras"**
|
||||
LazyVim provides "Packs" for specific needs. You can enable them in `lua/config/lazy.lua`:
|
||||
```lua
|
||||
require("lazyvim.util").plugin.setup({
|
||||
spec = {
|
||||
{ "LazyVim/LazyVim", import = "lazyvim.plugins" },
|
||||
-- Enable extras like Python, Docker, or Copilot:
|
||||
{ import = "lazyvim.plugins.extras.lang.python" },
|
||||
{ import = "lazyvim.plugins.extras.ui.mini-animate" },
|
||||
{ import = "lazyvim.plugins.extras.coding.copilot" },
|
||||
{ import = "lua.plugins" },
|
||||
},
|
||||
})
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. Why LazyVim?
|
||||
Unlike building a config from scratch (which can take months to perfect) or using a "thick" distro like LunarVim (which can feel bloated), LazyVim gives you a professional-grade starting point that feels like **your** config. It provides the "glue" that makes LSP, completion, and UI tools work together seamlessly.
|
||||
|
||||
---
|
||||
**Next Steps:**
|
||||
- Run `:LazyHealth` to check your environment.
|
||||
- Install `lazygit` on your system to enable the `<leader>gg` shortcut.
|
||||
- Explore the [Official LazyVim Docs](https://www.lazyvim.org).
|
||||
@@ -0,0 +1,82 @@
|
||||
---
|
||||
tags: [neovim, treesitter, mason, telescope, coding-tools, tutorial]
|
||||
created: 2026-05-29
|
||||
status: complete
|
||||
type: lesson
|
||||
---
|
||||
|
||||
# Deep Dive: The Power Trio (Tree-sitter, Mason, & Telescope)
|
||||
|
||||
While LazyVim provides the framework, its "superpowers" come from three specific plugins that redefine how you interact with code. Understanding these tools in-depth will allow you to master your editor.
|
||||
|
||||
---
|
||||
|
||||
## 1. Tree-sitter: The Syntactic Engine
|
||||
Traditional editors use "Regex" (Regular Expressions) to highlight code. Regex is just fancy pattern matching. **Tree-sitter** is different: it is a **parser**. It builds a concrete syntax tree (CST) of your source file.
|
||||
|
||||
### Why it matters:
|
||||
- **Semantic Awareness:** It knows the difference between a variable, a function, and a type, even if they have the same name.
|
||||
- **Incremental Parsing:** It updates the tree instantly as you type, making it incredibly fast.
|
||||
- **Language-Agnostic:** One engine powers hundreds of languages.
|
||||
|
||||
### Key Modules:
|
||||
- **Highlighting:** Precise, context-aware colors.
|
||||
- **Incremental Selection:** Logical selection (select word -> select expression -> select function).
|
||||
- **Indentation:** Uses the tree structure to know exactly where a line should start.
|
||||
|
||||
### Pro Commands:
|
||||
- `:TSInstall <lang>`: Install a specific language parser.
|
||||
- `:InspectTree`: Opens a side window showing the actual tree structure of your code.
|
||||
- `:EditQuery`: Advanced tool to create custom highlights or behavior.
|
||||
|
||||
---
|
||||
|
||||
## 2. Mason.nvim: The Tool Manager
|
||||
In the past, installing an LSP (Language Server) or a Formatter was a nightmare. You had to use `npm`, `pip`, `go install`, and `cargo` manually. **Mason** abstracts all of this.
|
||||
|
||||
### The Architecture:
|
||||
- **Isolated Binaries:** Mason installs tools in `~/.local/share/nvim/mason/bin/`. It does **not** pollute your system global PATH.
|
||||
- **The "Bridge":** Mason only *downloads* the tools. Plugins like `mason-lspconfig` are needed to connect those downloads to Neovim's internal LSP client.
|
||||
|
||||
### Essential Workflow:
|
||||
1. Open `:Mason`.
|
||||
2. Find your tool (e.g., `ruff` for Python, `gopls` for Go).
|
||||
3. Press `i` to install.
|
||||
4. Use `ensure_installed` in your config to automate this across machines.
|
||||
|
||||
---
|
||||
|
||||
## 3. Telescope.nvim: The Central Nervous System
|
||||
Telescope is a highly extendable fuzzy finder. It is the primary way you find and jump to information.
|
||||
|
||||
### The "Picker" Concept:
|
||||
Everything in Telescope is a **Picker**. A picker consists of three parts:
|
||||
1. **Finder:** Where the data comes from (files, git, buffers).
|
||||
2. **Sorter:** How it ranks the results (FZF, native).
|
||||
3. **Previewer:** Shows you the content before you click.
|
||||
|
||||
### Advanced Mappings (Inside Telescope):
|
||||
- `<C-j>` / `<C-k>`: Move through results.
|
||||
- `<Tab>`: Select multiple items.
|
||||
- `<C-q>`: Send all selected items to the **Quickfix List** (powerful for mass refactoring).
|
||||
- `<C-/>`: Show all keybindings for the current picker.
|
||||
|
||||
### Must-Have Extensions:
|
||||
- **`fzf-native`:** Uses a C-port of FZF for blazing fast sorting.
|
||||
- **`ui-select`:** Makes *all* Neovim selection menus (like Code Actions) use the Telescope UI.
|
||||
- **`undo`:** A visual search for your file's history.
|
||||
|
||||
---
|
||||
|
||||
## How They Work Together
|
||||
Imagine you are editing a file:
|
||||
1. **Mason** has installed the `Pyright` LSP server.
|
||||
2. **Tree-sitter** is providing beautiful syntax highlighting and logical selection.
|
||||
3. You realize you need to find where a function is used. You trigger **Telescope** (`lsp_references`).
|
||||
4. Telescope uses the LSP (installed by **Mason**) to find the references and displays them in its UI.
|
||||
|
||||
---
|
||||
**Next Steps:**
|
||||
- Try `:InspectTree` on a complex file to see how Tree-sitter "sees" your code.
|
||||
- Run `:Mason` and check if you have any outdated tools.
|
||||
- Experiment with `live_grep` in Telescope to find text across your entire project.
|
||||
@@ -0,0 +1,20 @@
|
||||
---
|
||||
title: Python Notes Index
|
||||
created: 2026-05-22
|
||||
tags:
|
||||
- python
|
||||
- index
|
||||
- moc
|
||||
category: index
|
||||
---
|
||||
# Python Notes Index
|
||||
|
||||
All notes under `note/python/`, grouped by subfolder.
|
||||
|
||||
```dataview
|
||||
LIST rows.file.link
|
||||
FROM "note"
|
||||
WHERE file.name != "index"
|
||||
GROUP BY file.folder
|
||||
SORT file.folder ASC
|
||||
```
|
||||
@@ -0,0 +1,176 @@
|
||||
# SQLAlchemy 2.0 Tutorial (Recipe Scraper Edition)
|
||||
|
||||
This guide covers SQLAlchemy 2.0, the industry-standard SQL toolkit and Object-Relational Mapper (ORM) for Python. We'll use the **Recipe Web Scraper** project models as our primary examples.
|
||||
|
||||
---
|
||||
|
||||
## 1. What is SQLAlchemy?
|
||||
|
||||
SQLAlchemy has two main components:
|
||||
1. **Core**: A SQL abstraction layer (SQL Expression Language, Schema definitions, Engine).
|
||||
2. **ORM**: A layer on top of Core that maps Python classes to database tables.
|
||||
|
||||
In this project, we primarily use the **ORM** to treat recipes and ingredients as Python objects.
|
||||
|
||||
---
|
||||
|
||||
## 2. Defining Models (The Modern Way)
|
||||
|
||||
SQLAlchemy 2.0 introduced a type-hint-centric way to define models using `Mapped` and `mapped_column`.
|
||||
|
||||
### The Base Class
|
||||
All models inherit from a common `Base` class created from `DeclarativeBase`.
|
||||
|
||||
```python
|
||||
from sqlalchemy.orm import DeclarativeBase
|
||||
|
||||
class Base(DeclarativeBase):
|
||||
pass
|
||||
```
|
||||
|
||||
### Example: The Recipe Model
|
||||
```python
|
||||
from sqlalchemy import String, Integer, DateTime
|
||||
from sqlalchemy.orm import Mapped, mapped_column, relationship
|
||||
from datetime import datetime
|
||||
|
||||
class Recipe(Base):
|
||||
__tablename__ = "recipes" # Name of the table in the DB
|
||||
|
||||
# Primary Key
|
||||
id: Mapped[int] = mapped_column(primary_key=True)
|
||||
|
||||
# Simple Columns (SQLAlchemy infers types from Mapped[T])
|
||||
url: Mapped[str] = mapped_column(String, unique=True, index=True)
|
||||
title: Mapped[str]
|
||||
total_time: Mapped[int | None] # Optional column (nullable=True)
|
||||
|
||||
# Column with a default value
|
||||
scraped_at: Mapped[datetime] = mapped_column(DateTime, default=datetime.utcnow)
|
||||
|
||||
# Relationships (Defined in section 4)
|
||||
ingredients: Mapped[list["Ingredient"]] = relationship(back_populates="recipe")
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Engine and Session
|
||||
|
||||
### The Engine
|
||||
The **Engine** is the starting point for any SQLAlchemy application. It manages a pool of connections to the database.
|
||||
|
||||
```python
|
||||
from sqlalchemy import create_engine
|
||||
|
||||
# SQLite: The '///' means relative path to the current directory
|
||||
engine = create_engine("sqlite:///recipes.db", echo=True)
|
||||
# echo=True logs all SQL commands to the terminal (great for debugging)
|
||||
```
|
||||
|
||||
### Creating Tables
|
||||
You can tell SQLAlchemy to create all tables defined in your models:
|
||||
```python
|
||||
Base.metadata.create_all(engine)
|
||||
```
|
||||
|
||||
### The Session
|
||||
The **Session** handles the conversation with the database. Use `sessionmaker` to create a factory for sessions.
|
||||
|
||||
```python
|
||||
from sqlalchemy.orm import sessionmaker
|
||||
|
||||
SessionLocal = sessionmaker(bind=engine)
|
||||
|
||||
# Use as a context manager to ensure the connection is closed
|
||||
with SessionLocal() as session:
|
||||
# do work here
|
||||
pass
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. Relationships (1-to-Many)
|
||||
|
||||
In our project, one `Recipe` has many `Ingredients`.
|
||||
|
||||
### Foreign Key
|
||||
The "child" table (`Ingredient`) must have a column pointing to the "parent" table (`Recipe`).
|
||||
|
||||
```python
|
||||
from sqlalchemy import ForeignKey
|
||||
|
||||
class Ingredient(Base):
|
||||
__tablename__ = "ingredients"
|
||||
id: Mapped[int] = mapped_column(primary_key=True)
|
||||
|
||||
# Links to 'recipes.id'
|
||||
recipe_id: Mapped[int] = mapped_column(ForeignKey("recipes.id"))
|
||||
|
||||
text: Mapped[str]
|
||||
|
||||
# Back-reference to the parent Recipe object
|
||||
recipe: Mapped["Recipe"] = relationship(back_populates="ingredients")
|
||||
```
|
||||
|
||||
### Cascades
|
||||
`cascade="all, delete-orphan"` ensures that if you delete a Recipe, all its Ingredients are also deleted automatically.
|
||||
|
||||
---
|
||||
|
||||
## 5. CRUD Operations
|
||||
|
||||
### Create (Insert)
|
||||
```python
|
||||
with SessionLocal() as session:
|
||||
new_recipe = Recipe(title="Pasta Carbonara", url="https://example.com/pasta")
|
||||
session.add(new_recipe)
|
||||
session.commit() # Save to DB
|
||||
```
|
||||
|
||||
### Read (Select)
|
||||
```python
|
||||
from sqlalchemy import select
|
||||
|
||||
with SessionLocal() as session:
|
||||
# 1. Get by ID
|
||||
recipe = session.get(Recipe, 1)
|
||||
|
||||
# 2. Filter by column
|
||||
stmt = select(Recipe).where(Recipe.title == "Pasta Carbonara")
|
||||
result = session.execute(stmt).scalars().first()
|
||||
|
||||
# 3. Get all
|
||||
all_recipes = session.query(Recipe).all() # Older syntax, still common
|
||||
```
|
||||
|
||||
### Update
|
||||
```python
|
||||
with SessionLocal() as session:
|
||||
recipe = session.get(Recipe, 1)
|
||||
recipe.title = "Authentic Pasta Carbonara"
|
||||
session.commit()
|
||||
```
|
||||
|
||||
### Delete
|
||||
```python
|
||||
with SessionLocal() as session:
|
||||
recipe = session.get(Recipe, 1)
|
||||
session.delete(recipe)
|
||||
session.commit()
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. Common Pitfalls
|
||||
|
||||
1. **Lazy Loading**: By default, SQLAlchemy doesn't load relationships until you access them. This can cause "N+1" performance issues. Use `joinedload` to fetch everything in one query.
|
||||
2. **Session Lifecycle**: Always use a context manager (`with session:`) or close your sessions manually.
|
||||
3. **Commit vs Flush**: `session.flush()` sends changes to the DB but doesn't permanentize them. `session.commit()` makes them permanent.
|
||||
|
||||
---
|
||||
|
||||
## 7. Next Steps: Migrations with Alembic
|
||||
|
||||
As your models change (e.g., you add a `rating` column), you shouldn't just delete the DB and start over. **Alembic** is the tool used to handle database migrations.
|
||||
|
||||
Install it with: `pip install alembic`
|
||||
@@ -0,0 +1,93 @@
|
||||
---
|
||||
title: Poetry Setup and Project Initialization Guide
|
||||
created: 2026-05-22
|
||||
tags:
|
||||
- python
|
||||
- python_tool
|
||||
- poetry
|
||||
- guide
|
||||
category: python_tool
|
||||
status: reference
|
||||
up: "[[project/index]]"
|
||||
related:
|
||||
- "[[poetry_project_ideas]]"
|
||||
- "[[__init__.py explained]]"
|
||||
source:
|
||||
author:
|
||||
published:
|
||||
---
|
||||
# Poetry Setup and Project Initialization Guide
|
||||
|
||||
This guide explains how to install Poetry and start a new Python project, based on the concepts from "Introduction to Poetry - Python Dependency Management".
|
||||
|
||||
## 1. How to Install Poetry
|
||||
|
||||
While the introductory notes focus on usage, the standard way to install Poetry is via the official installer script.
|
||||
|
||||
### macOS / Linux / WSL
|
||||
Open your terminal and run:
|
||||
```bash
|
||||
curl -sSL https://install.python-poetry.org | python3 -
|
||||
```
|
||||
|
||||
### Windows (PowerShell)
|
||||
Open PowerShell and run:
|
||||
```powershell
|
||||
(Invoke-WebRequest -Uri https://install.python-poetry.org -UseBasicParsing).Content | py -
|
||||
```
|
||||
|
||||
### Verification
|
||||
After installation, restart your terminal and verify by running:
|
||||
```bash
|
||||
poetry --version
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. How to Start a Project
|
||||
|
||||
There are two main ways to start a project with Poetry:
|
||||
|
||||
### Method A: Creating a New Project (Recommended for new folders)
|
||||
To create a new project with a predefined folder structure:
|
||||
```bash
|
||||
poetry new my-project
|
||||
```
|
||||
This creates a directory named `my-project` with the following structure:
|
||||
```text
|
||||
my-project/
|
||||
├── pyproject.toml
|
||||
├── README.md
|
||||
├── my_project/
|
||||
│ └── __init__.py
|
||||
└── tests/
|
||||
└── __init__.py
|
||||
```
|
||||
|
||||
### Method B: Initializing an Existing Project
|
||||
If you already have a project folder and want to add Poetry to it:
|
||||
1. Navigate to your project directory:
|
||||
```bash
|
||||
cd my-existing-project
|
||||
```
|
||||
2. Run the interactive initialization command mentioned in the introduction:
|
||||
```bash
|
||||
poetry init
|
||||
```
|
||||
This will walk you through creating your `pyproject.toml` file interactively.
|
||||
|
||||
---
|
||||
|
||||
## 3. Key Concepts from the Introduction
|
||||
|
||||
- **`pyproject.toml`**: The single source of truth for your project configuration (replaces `requirements.txt`, `setup.py`, etc.).
|
||||
- **Deterministic Resolution**: Poetry ensures your dependencies are resolved correctly using a lockfile (`poetry.lock`).
|
||||
- **Isolation**: Poetry automatically manages virtual environments for you, ensuring your global Python installation stays clean.
|
||||
|
||||
## 4. Basic Workflow Commands
|
||||
|
||||
Once your project is started, use these commands to manage it:
|
||||
- `poetry add <package>`: Add and install a new dependency.
|
||||
- `poetry install`: Install all dependencies defined in `pyproject.toml`.
|
||||
- `poetry shell`: Activate the project's virtual environment.
|
||||
- `poetry run <command>`: Run a command inside the virtual environment without activating it.
|
||||
@@ -0,0 +1,152 @@
|
||||
---
|
||||
title: What is __init__.py?
|
||||
created: 2026-05-22
|
||||
tags:
|
||||
- python
|
||||
- python_note
|
||||
- packaging
|
||||
- reference
|
||||
category: python_note
|
||||
status: reference
|
||||
up: "[[project/index]]"
|
||||
related:
|
||||
- "[[poetry_guide]]"
|
||||
- "[[poetry_project_ideas]]"
|
||||
source:
|
||||
author:
|
||||
published:
|
||||
---
|
||||
# What is `__init__.py`?
|
||||
|
||||
`__init__.py` is the file that tells Python "this folder is a package". When you put it inside a directory, that directory becomes importable like a module.
|
||||
|
||||
## The Core Purpose
|
||||
|
||||
Without `__init__.py` (in older Python, ≤3.2), a folder was just a folder — Python could not `import` from it. With it, the folder becomes a **package** you can do this with:
|
||||
|
||||
```python
|
||||
from recipe_scraper.scraper import fetch_recipe
|
||||
```
|
||||
|
||||
Here, `recipe_scraper/` is a package because it contains `__init__.py`.
|
||||
|
||||
> Note: Since Python 3.3, "namespace packages" allow imports without `__init__.py`, but **regular packages still use it** because it gives you more control (init code, explicit exports, IDE/tooling support).
|
||||
|
||||
---
|
||||
|
||||
## What It Does in Practice
|
||||
|
||||
### 1. Marks the folder as a package
|
||||
Even an **empty** `__init__.py` is meaningful. It signals to Python:
|
||||
> "Treat this directory as something you can import from."
|
||||
|
||||
### 2. Runs initialization code
|
||||
Anything inside `__init__.py` runs **the first time** the package is imported. This is useful for:
|
||||
- Setting up logging
|
||||
- Loading config
|
||||
- Registering plugins
|
||||
|
||||
```python
|
||||
# recipe_scraper/__init__.py
|
||||
import logging
|
||||
logging.getLogger(__name__).addHandler(logging.NullHandler())
|
||||
```
|
||||
|
||||
### 3. Controls the package's public API
|
||||
You can re-export things so users don't need to know the internal file layout:
|
||||
|
||||
```python
|
||||
# recipe_scraper/__init__.py
|
||||
from .scraper import fetch_recipe
|
||||
from .parser import parse_recipe
|
||||
from .db import save_recipe
|
||||
|
||||
__all__ = ["fetch_recipe", "parse_recipe", "save_recipe"]
|
||||
```
|
||||
|
||||
Now consumers can write the short form:
|
||||
```python
|
||||
from recipe_scraper import fetch_recipe # clean
|
||||
# instead of:
|
||||
from recipe_scraper.scraper import fetch_recipe # verbose
|
||||
```
|
||||
|
||||
### 4. Defines package metadata
|
||||
A common pattern is exposing a version string:
|
||||
```python
|
||||
# recipe_scraper/__init__.py
|
||||
__version__ = "0.1.0"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## How It Fits the Recipe Scraper Project
|
||||
|
||||
Given the folder structure from [[Recipe Web Scraper - Getting Started]]:
|
||||
|
||||
```
|
||||
recipe-scraper/
|
||||
├── pyproject.toml
|
||||
└── recipe_scraper/
|
||||
├── __init__.py ← makes this a package
|
||||
├── scraper.py
|
||||
├── parser.py
|
||||
├── db.py
|
||||
└── cli.py
|
||||
```
|
||||
|
||||
The `__init__.py` here lets you:
|
||||
|
||||
1. Run `poetry run python -m recipe_scraper.cli` — only works because `recipe_scraper` is a package.
|
||||
2. Import cleanly from anywhere in the project:
|
||||
```python
|
||||
from recipe_scraper.db import Recipe
|
||||
```
|
||||
3. Optionally expose a tidy top-level API:
|
||||
```python
|
||||
# recipe_scraper/__init__.py
|
||||
from .scraper import scrape_url
|
||||
from .db import Recipe, init_db
|
||||
|
||||
__version__ = "0.1.0"
|
||||
__all__ = ["scrape_url", "Recipe", "init_db"]
|
||||
```
|
||||
So users of your library can just do:
|
||||
```python
|
||||
import recipe_scraper
|
||||
recipe_scraper.scrape_url("https://...")
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## When to Leave It Empty vs. Fill It
|
||||
|
||||
| Situation | What to put in `__init__.py` |
|
||||
| :--- | :--- |
|
||||
| Internal-only package, no public API | Empty file |
|
||||
| You want a clean import surface | Re-exports + `__all__` |
|
||||
| Library shipped to PyPI | `__version__`, re-exports, maybe logging setup |
|
||||
| One-time setup needed (config, env) | Initialization code at the top |
|
||||
|
||||
**Rule of thumb:** start with an empty `__init__.py`. Only add code when you have a concrete reason — re-exports, version, or setup. Don't put heavy logic in `__init__.py`; it runs on every import.
|
||||
|
||||
---
|
||||
|
||||
## Common Gotchas
|
||||
|
||||
- **Heavy imports slow everything down.** If `__init__.py` imports a big library (e.g., `pandas`), every `import recipe_scraper.anything` pays that cost. Keep it light.
|
||||
- **Circular imports** often start in `__init__.py`. If `__init__.py` imports from `scraper.py`, and `scraper.py` imports from the package root, you get a cycle. Use lazy imports or restructure.
|
||||
- **Tests need it too.** A `tests/` folder usually has an empty `__init__.py` so pytest can discover test modules consistently (though pytest's `rootdir` config can avoid this).
|
||||
|
||||
---
|
||||
|
||||
## TL;DR
|
||||
|
||||
- `__init__.py` = "this folder is a Python package".
|
||||
- Can be empty — just its presence matters.
|
||||
- Use it to **re-export** the public API, set `__version__`, or run small setup.
|
||||
- Keep it light: every import of the package runs it.
|
||||
|
||||
## Related
|
||||
- [[Recipe Web Scraper - Getting Started]]
|
||||
- [[poetry_guide]] — Poetry's `poetry new` creates this file automatically
|
||||
@@ -0,0 +1,177 @@
|
||||
---
|
||||
title: __init__.py vs main.py
|
||||
created: 2026-05-22
|
||||
tags:
|
||||
- python
|
||||
- python_note
|
||||
- packaging
|
||||
- reference
|
||||
category: python_note
|
||||
status: reference
|
||||
up: "[[project/index]]"
|
||||
related:
|
||||
- "[[__init__.py explained]]"
|
||||
- "[[Recipe Web Scraper - Getting Started]]"
|
||||
source:
|
||||
author:
|
||||
published:
|
||||
---
|
||||
# `__init__.py` vs `main.py`
|
||||
|
||||
Both files live inside a Python package, but they have very different jobs. This note uses the [[Recipe Web Scraper - Getting Started|recipe scraper]] project as the running example.
|
||||
|
||||
## The Recipe Scraper Layout
|
||||
|
||||
```
|
||||
recipe-scraper/
|
||||
└── src/
|
||||
└── recipe_scraper/
|
||||
├── __init__.py ← marks this as a package
|
||||
├── main.py ← CLI entry point (you run this)
|
||||
├── scraper.py
|
||||
├── models.py
|
||||
├── db.py
|
||||
└── exporters.py
|
||||
```
|
||||
|
||||
Two files, two completely different purposes.
|
||||
|
||||
---
|
||||
|
||||
## At a Glance
|
||||
|
||||
| Aspect | `__init__.py` | `main.py` |
|
||||
| :--- | :--- | :--- |
|
||||
| **Purpose** | Marks the folder as a package; sets up the package | Holds the **program's entry point** (the code you run) |
|
||||
| **Required?** | Yes (for regular packages) | No — it's a convention, name can be anything |
|
||||
| **When it runs** | Automatically, on **any** import of the package | Only when you **explicitly run it** |
|
||||
| **Typical contents** | Re-exports, `__version__`, light setup | `main()` function, CLI parsing, `if __name__ == "__main__"` |
|
||||
| **Who calls it** | Python's import system | The user (via `python -m`, `poetry run`, etc.) |
|
||||
| **Should be lightweight?** | **Yes** — runs every import | Can be heavy — runs once when invoked |
|
||||
|
||||
---
|
||||
|
||||
## `__init__.py` — The Package Setup File
|
||||
|
||||
Its job is to tell Python "this folder is a package" and optionally configure what the package looks like from the outside.
|
||||
|
||||
```python
|
||||
# src/recipe_scraper/__init__.py
|
||||
from .scraper import scrape_url
|
||||
from .db import Recipe, init_db
|
||||
|
||||
__version__ = "0.1.0"
|
||||
__all__ = ["scrape_url", "Recipe", "init_db"]
|
||||
```
|
||||
|
||||
After this, anyone using the library can write:
|
||||
```python
|
||||
from recipe_scraper import scrape_url
|
||||
scrape_url("https://example.com/recipe")
|
||||
```
|
||||
|
||||
instead of digging into `recipe_scraper.scraper`. **It never gets "run" directly — it runs implicitly whenever the package is imported.**
|
||||
|
||||
See [[__init__.py explained]] for the full picture.
|
||||
|
||||
---
|
||||
|
||||
## `main.py` — The Program Entry Point
|
||||
|
||||
Its job is to be the **code that actually executes when you launch the program**. For the recipe scraper, it's where the CLI lives:
|
||||
|
||||
```python
|
||||
# src/recipe_scraper/main.py
|
||||
import typer
|
||||
from .scraper import scrape_url
|
||||
from .db import init_db
|
||||
|
||||
app = typer.Typer()
|
||||
|
||||
@app.command()
|
||||
def scrape(url: str):
|
||||
"""Scrape a recipe URL and save it to the database."""
|
||||
init_db()
|
||||
recipe = scrape_url(url)
|
||||
print(f"Saved: {recipe.title}")
|
||||
|
||||
@app.command()
|
||||
def list_recipes():
|
||||
"""List all saved recipes."""
|
||||
...
|
||||
|
||||
def main():
|
||||
app()
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
```
|
||||
|
||||
You run it like:
|
||||
```bash
|
||||
poetry run python -m recipe_scraper.main scrape https://example.com/recipe
|
||||
```
|
||||
|
||||
Or, if `pyproject.toml` defines an entry point pointing at `main:main`, just:
|
||||
```bash
|
||||
poetry run recipe-scraper scrape https://example.com/recipe
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## The Key Mental Model
|
||||
|
||||
> `__init__.py` answers: **"What is this package?"**
|
||||
> `main.py` answers: **"What happens when you run this program?"**
|
||||
|
||||
- Importing the package → `__init__.py` runs.
|
||||
- Running the program → `main.py` runs (and it imports things, which triggers `__init__.py`).
|
||||
|
||||
So in a typical execution:
|
||||
```
|
||||
$ poetry run python -m recipe_scraper.main scrape ...
|
||||
│
|
||||
├── Python loads `recipe_scraper` package
|
||||
│ └── runs __init__.py (sets up exports, version, etc.)
|
||||
│
|
||||
└── Python runs main.py as the module
|
||||
└── parses CLI args, calls scrape_url(), saves to DB
|
||||
```
|
||||
|
||||
`__init__.py` runs **first** (because `main.py` is inside the package), but it does almost nothing visible. `main.py` is where the actual work happens.
|
||||
|
||||
---
|
||||
|
||||
## Common Pitfalls
|
||||
|
||||
### 1. Putting CLI code in `__init__.py`
|
||||
**Don't.** It will run every time *anything* imports your package, including your tests. Keep entry-point code in `main.py` (or `cli.py`).
|
||||
|
||||
### 2. Heavy imports in `__init__.py`
|
||||
If `__init__.py` imports `pandas`, `torch`, or anything slow, every import of your package pays that cost. For the recipe scraper, avoid importing `recipe-scrapers` or SQLAlchemy at the top of `__init__.py` — import them inside the modules that need them.
|
||||
|
||||
### 3. Forgetting `if __name__ == "__main__":` in `main.py`
|
||||
Without that guard, `main.py`'s code runs whenever the module is **imported**, not just when it's **executed**. Always wrap the entry-point call:
|
||||
```python
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
```
|
||||
|
||||
### 4. Confusing `main.py` with `__main__.py`
|
||||
- `main.py` — just a convention; you have to point at it explicitly (`python -m recipe_scraper.main`).
|
||||
- `__main__.py` — a **special name**: lets you run `python -m recipe_scraper` (no `.main` needed).
|
||||
|
||||
If you want `python -m recipe_scraper` to work, rename `main.py` to `__main__.py`. Many projects keep both: `__main__.py` is a one-liner that calls into `main.py`.
|
||||
|
||||
---
|
||||
|
||||
## TL;DR
|
||||
|
||||
- **`__init__.py`** = package marker + public-API shaping. Runs on every import. Keep it light.
|
||||
- **`main.py`** = program entry point. Runs when you invoke the CLI. This is where the action is.
|
||||
- They're complementary, not alternatives. In the recipe scraper, `__init__.py` exposes the library; `main.py` is the CLI you actually run.
|
||||
|
||||
## Related
|
||||
- [[__init__.py explained]]
|
||||
- [[Recipe Web Scraper - Getting Started]]
|
||||
- [[poetry_guide]]
|
||||
Reference in New Issue
Block a user