vault backup: 2026-05-30 16:02:33

This commit is contained in:
Rainyy21
2026-05-30 16:02:33 -04:00
commit 96449f8968
43 changed files with 2837 additions and 0 deletions
+2
View File
@@ -0,0 +1,2 @@
/../project/recipe web scraper/respect robots.txt check with urlib.robotparser.md
+10
View File
@@ -0,0 +1,10 @@
---
tags:
- coding
---
```dataview
TABLE file.tags AS Tags, file.mtime AS Modified
FROM "devop_note"
WHERE file.name != "00_devop_note"
SORT file.name ASC
```
+6
View File
@@ -0,0 +1,6 @@
---
aliases:
up: "[[00_devop_note]]"
tags:
- devop
---
+5
View File
@@ -0,0 +1,5 @@
---
tags:
- devop
up: "[[00_devop_note]]"
---
+8
View File
@@ -0,0 +1,8 @@
---
tags:
- devop
up: "[[00_devop_note]]"
---
# What is Maven?
maven is a build automation tool used primarily for java projects, hosted by Apache Software Foundation. Maven projects are configured using Project Object Model (POM) in a `pom.xml` file. It's used by over 70% of Java organizations, so employers actively seek people with strong Maven skills
+175
View File
@@ -0,0 +1,175 @@
---
aliases:
up: "[[00_devop_note]]"
tags:
- devop
---
# Uvicorn and Gunicorn
Both are Python application servers — they run your Python web app and hand requests back and forth with a web server like Nginx. But they solve **different problems**, and you'll often see them used **together**.
The key split: **Gunicorn is WSGI** (synchronous), **Uvicorn is ASGI** (asynchronous).
## Quick refresher: WSGI vs ASGI
| | WSGI | ASGI |
|---|---|---|
| Spec | PEP 3333 | "Asynchronous Server Gateway Interface" |
| Model | Sync, one request per worker at a time | Async, many concurrent requests per worker |
| Frameworks | Django (classic), Flask | FastAPI, Starlette, Django (async views), Sanic |
| Supports WebSockets? | No | Yes |
| Supports HTTP/2, SSE? | No | Yes |
If your app uses `async def` view functions or WebSockets, you need ASGI.
If it's a traditional Django/Flask app with `def` views, WSGI is fine.
---
## Gunicorn ("Green Unicorn")
A **WSGI** server. Mature, simple, battle-tested. The de facto default for Django/Flask in production.
### Why people pick it
- **Simple config** — most options are sensible by default
- **Pre-fork worker model** — master process forks N workers; each worker handles one request at a time
- **Stable** — has been the standard for years
- **Good signal handling** — graceful reloads, zero-downtime restarts
### Minimal usage
```bash
gunicorn myproject.wsgi:application --workers 4 --bind 0.0.0.0:8000
```
Or with a config file `gunicorn.conf.py`:
```python
bind = "unix:/tmp/myproject.sock"
workers = 4
worker_class = "sync" # default
timeout = 30
accesslog = "-"
errorlog = "-"
```
### Worker classes
Gunicorn lets you swap the worker type:
- `sync` — default, one request at a time per worker
- `gthread` — threaded workers (good for I/O-bound apps)
- `gevent` / `eventlet` — async via greenlets (legacy)
- `uvicorn.workers.UvicornWorker` — **this is the bridge to ASGI** (see below)
### How many workers?
Rule of thumb: `(2 × CPU cores) + 1`. So a 4-core box → 9 workers.
---
## Uvicorn
An **ASGI** server built on `uvloop` and `httptools` — both written in C, which makes it very fast. It's the standard server for FastAPI and modern async Python web apps.
### Why people pick it
- **Async-native** — handles thousands of concurrent connections per worker
- **WebSockets + HTTP/2 support**
- **Very fast** — uvloop is a drop-in faster replacement for asyncio's event loop
- **Lightweight** — small dependency surface
### Minimal usage
```bash
uvicorn myproject.main:app --host 0.0.0.0 --port 8000
```
For development with auto-reload:
```bash
uvicorn myproject.main:app --reload
```
### Where it falls short alone
Uvicorn by itself is **single-process**. To use multiple CPU cores in production, you need to either:
1. Run multiple Uvicorn instances behind a load balancer, OR
2. Run Uvicorn **inside Gunicorn** as worker processes (the common pattern)
---
## The common production combo: Gunicorn + Uvicorn
This is the standard FastAPI production setup:
```bash
gunicorn myproject.main:app \
--workers 4 \
--worker-class uvicorn.workers.UvicornWorker \
--bind 0.0.0.0:8000
```
What's happening:
- **Gunicorn** is the process manager — forks workers, handles signals, restarts dead workers, manages graceful shutdowns
- **Uvicorn** runs *inside* each Gunicorn worker — provides the ASGI event loop that runs your async code
You get the best of both: Gunicorn's robust process management + Uvicorn's async performance.
### Visual
```
┌─ Uvicorn worker (async loop) ─ FastAPI app
│
Nginx → Gunicorn ├─ Uvicorn worker (async loop) ─ FastAPI app
(master) │
├─ Uvicorn worker (async loop) ─ FastAPI app
│
└─ Uvicorn worker (async loop) ─ FastAPI app
```
---
## Decision guide
| Your situation | Use |
|---|---|
| Django (sync), Flask | **Gunicorn** alone |
| FastAPI, Starlette, async Django | **Gunicorn + UvicornWorker** |
| Local dev, single FastAPI process | **Uvicorn** alone (`--reload`) |
| WebSockets required | Must be ASGI → **Uvicorn** (alone or under Gunicorn) |
| Legacy app, uWSGI already configured | [[uWSGI]] — no urgent need to migrate |
---
## Comparison with [[uWSGI]]
| | Gunicorn | Uvicorn | uWSGI |
|---|---|---|---|
| Protocol | WSGI | ASGI | WSGI (+ many others) |
| Async support | No (sync workers) | Yes (native) | Limited |
| Config complexity | Low | Low | **Very high** |
| WebSockets | No | Yes | Partial |
| Speed (raw) | Good | **Fastest** for async | Fast but heavy |
| Maintained actively | Yes | Yes | Concerns |
| Best for | Django/Flask | FastAPI | Legacy / specialized features |
---
## Common gotchas
1. **Don't run Uvicorn `--reload` in production** — it's a dev-only feature, has overhead and isn't safe.
2. **Workers ≠ threads** — each Gunicorn worker is a separate Python process with its own memory. Database connections, in-memory caches, etc. are *not* shared between workers.
3. **Timeouts matter** — Gunicorn's default `timeout=30s` will kill workers running long async tasks. Tune it for your workload.
4. **Nginx is still recommended in front** — Uvicorn/Gunicorn don't do SSL termination, static files, or rate limiting as well as Nginx.
5. **Logging** — by default both log to stdout/stderr; route to your log aggregator via container stdout (`-` in config).
---
## Related
- [[uWSGI]] — older alternative, mostly WSGI
- [[Apache Tomcat]] — the Java equivalent (servlet container)
+94
View File
@@ -0,0 +1,94 @@
---
aliases:
up: "[[00_devop_note]]"
tags:
- devop
---
w
# What is uWSGI?
**uWSGI** is an application server that sits between a web server (like Nginx or Apache) and a Python web application (like Django, Flask, or FastAPI). It runs your Python code and handles incoming web requests.
The name comes from **WSGI** (Web Server Gateway Interface), which is the standard Python spec (PEP 3333) that defines how web servers talk to Python applications. The lowercase "u" is the Greek letter μ (micro), suggesting it's lightweight — though in practice it grew into a full-featured server.
## Why do we need it?
A plain web server like Nginx doesn't know how to execute Python code. It only serves static files (HTML, CSS, images) and forwards dynamic requests elsewhere. You need a process that:
1. Loads your Python application into memory
2. Receives requests from the web server
3. Calls your Python code with the request
4. Returns the response back to the web server
That middle process is uWSGI (or alternatives like Gunicorn).
## The typical stack
```
Browser → Nginx → uWSGI → Python app (Django/Flask)
```
- **Nginx**: handles SSL, static files, load balancing, gzip, caching
- **uWSGI**: runs Python workers, manages processes/threads
- **Python app**: your business logic
Nginx and uWSGI usually talk over a **Unix socket** (fast, local) or a TCP port. The protocol between them is called the **uwsgi protocol** (lowercase) — a binary protocol that's faster than plain HTTP.
## Key features
- **Process management**: spawns multiple worker processes to handle concurrent requests
- **Threading**: each worker can run multiple threads
- **Auto-reload**: restart workers when code changes (dev mode)
- **Emperor mode**: one master process supervising many vassal apps
- **Cheaper mode**: dynamically scale workers up/down based on load
- **Multi-language**: despite the name, also supports Ruby, Perl, Go, etc.
## Minimal config example
A typical `uwsgi.ini`:
```ini
[uwsgi]
module = myproject.wsgi:application
master = true
processes = 4
threads = 2
socket = /tmp/myproject.sock
chmod-socket = 660
vacuum = true
die-on-term = true
```
- `module`: entry point (the WSGI callable)
- `processes`: how many worker processes to fork
- `socket`: where Nginx connects to
- `vacuum`: clean up the socket on exit
- `die-on-term`: shut down cleanly on SIGTERM
## Running it
```bash
uwsgi --ini uwsgi.ini
```
Or in production, run it under **systemd** so it restarts on failure.
## uWSGI vs Gunicorn
Both are WSGI servers. The community has largely shifted toward **Gunicorn** because:
- Simpler config
- Fewer footguns
- Easier to deploy
uWSGI is more feature-rich and faster in some benchmarks, but its config surface is huge (hundreds of options) and the project has had governance/maintenance concerns. For new projects, Gunicorn behind Nginx is the common default.
## When you'll see uWSGI
- Legacy Django/Flask deployments
- Setups that need uWSGI-specific features (Emperor, cheaper, etc.)
- Docker images for older Python web apps
## Related
- [[Apache Tomcat]] — the Java equivalent role (servlet container running Java apps behind a web server)
@@ -0,0 +1,55 @@
---
tags:
- devop
- video
- poetry
- python
up: "[[00_devop_note]]"
related:
- "[[Poetry]]"
- "[[poetry_guide]]"
- "[[poetry_project_ideas]]"
source:
author:
published:
---
# Poetry: Dependency Management for Python
Poetry is a tool for **dependency management** and **packaging** in Python. It allows you to declare the libraries your project depends on and it will manage (install/update) them for you. Poetry offers a lockfile to ensure repeatable installs, and can build your project for distribution.
## Similarities to Maven (Java)
If you are familiar with Apache Maven, you can think of Poetry as serving a similar purpose in the Python ecosystem:
| Feature | Maven (Java) | Poetry (Python) |
| :--- | :--- | :--- |
| **Project Configuration** | `pom.xml` | `pyproject.toml` |
| **Dependency Resolution** | Resolves transitive dependencies | Resolves transitive dependencies with a deterministic solver |
| **Locking** | No direct equivalent (uses version ranges in POM) | `poetry.lock` (ensures exact versions) |
| **Build & Packaging** | Builds JAR/WAR files | Builds Wheel and sdist packages |
| **Publishing** | Deploys to Central/Nexus | Publishes to PyPI or private repositories |
| **Environment Management**| Relies on external JRE/JDK | Manages Virtual Environments automatically |
## Key Concepts
### 1. `pyproject.toml`
This is the single source of truth for your project. It replaces `setup.py`, `requirements.txt`, `setup.cfg`, `MANIFEST.in` and `pipfile`.
### 2. Deterministic Resolution
Poetry comes with a custom dependency resolver that will always find a solution if one exists, or clearly explain why it failed.
### 3. Isolation by Default
Poetry always runs in isolation. It either uses your existing virtual environment or creates its own to ensure that your project dependencies don't leak into your global Python installation.
## Basic Commands
- `poetry init`: Interactively create a `pyproject.toml` file.
- `poetry add <package>`: Adds a dependency to `pyproject.toml` and installs it.
- `poetry install`: Installs the dependencies specified in `pyproject.toml` (or `poetry.lock` if present).
- `poetry update`: Updates dependencies to their latest versions according to `pyproject.toml` and updates the lock file.
- `poetry run <command>`: Runs a command within the project's virtual environment.
- `poetry shell`: Spawns a shell within the virtual environment.
## Why use Poetry over Pip?
While `pip` is the standard package installer, it doesn't handle dependency resolution or project metadata as comprehensively as Poetry. Poetry provides a more "all-in-one" experience, much like Maven does for Java, by combining dependency management, environment isolation, and packaging into a single tool.
@@ -0,0 +1,87 @@
---
tags:
- devop
- video
- maven
- java
up: "[[00_devop_note]]"
related: "[[Maven]]"
source: https://www.youtube.com/watch?v=T00NKLQvwYE
author: Cameron McKenzie (TheServerSide)
published: 2023-08-27
---
# Learn Apache Maven Full Tutorial in Java for Beginners
> An hour-long beginner-friendly course that walks through Apache Maven from scratch — installation, configuration, core commands, dependencies, plugins, and integrations with Jenkins and Docker.
## Overview
Apache Maven is the most widely used build automation tool in the Java ecosystem (used by ~70% of Java organizations). This tutorial progresses from foundational concepts (installing Java/JDK and Maven) through to advanced topics like building cloud-native microservices and CI/CD integration.
## Topics Covered
### 1. Getting Started
- What Maven is and what it enables developers to do
- Installing the JDK (prerequisite)
- Downloading, installing, and configuring Maven
- Verifying the install with `mvn -v`
### 2. Core Maven Commands
- `mvn compile` — compile source code
- `mvn test` — run unit tests
- `mvn package` — bundle compiled code into a JAR/WAR
- `mvn install` — install artifact into local repo
- `mvn deploy` — push artifact to a remote repo
- `mvn clean` — wipe the `target/` directory
### 3. The POM File (`pom.xml`)
- The heart of every Maven project
- POM types (parent, aggregator, effective POM)
- Properties and how they control build behavior
- Project coordinates: `groupId`, `artifactId`, `version`
### 4. Dependency Management
- Declaring dependencies
- **Scopes**: `compile`, `provided`, `runtime`, `test`, `system`
- External dependencies
- Exclusions and optional dependencies
- How Maven resolves transitive dependencies
### 5. Plugins
- Plugins extend Maven's core functionality
- Commonly used:
- **Compiler Plugin** — controls the Java source/target version
- **Surefire Plugin** — runs unit tests
- When to add a plugin vs. rely on defaults
### 6. Maven vs. Gradle
- Comparison of the two dominant JVM build tools
- Tradeoffs: XML configuration (Maven) vs. Groovy/Kotlin DSL (Gradle)
- Why Maven still dominates enterprise Java
### 7. IDE Integration
- Creating and managing Maven projects in **IntelliJ IDEA** and **Eclipse**
- Building projects from the IDE
- Creating executable JARs
- Handling multi-module projects
### 8. DevOps Integration
- **Maven + Jenkins** — running builds in CI pipelines
- **Maven + Docker** — packaging artifacts into Docker images
- Building cloud-native microservices
## Key Takeaways
- Maven is *opinionated* — follow the standard directory layout (`src/main/java`, `src/test/java`) and most things just work
- The `pom.xml` is declarative; you describe *what* the project is, not *how* to build it
- Dependency scopes matter — using `test` scope keeps test libraries out of production artifacts
- Plugins are how Maven gets extended; nearly every advanced behavior is plugin-driven
## Actionable Next Steps
- [ ] Install JDK and Maven; verify with `mvn -v`
- [ ] Generate a starter project: `mvn archetype:generate`
- [ ] Try the full lifecycle: `mvn clean package`
- [ ] Add a dependency from [Maven Central](https://search.maven.org)
- [ ] Wire a simple project into Jenkins or Docker
## Source
- Video: [Learn Apache Maven Full Tutorial in Java for Beginners](https://www.youtube.com/watch?v=T00NKLQvwYE)
- Companion article: [TheServerSide — Learn Maven tutorial for beginners](https://www.theserverside.com/video/Learn-Maven-tutorial-for-beginners)
+32
View File
@@ -0,0 +1,32 @@
# 🗂️ LeetCode Note Index
Welcome to your LeetCode problem index. This dashboard automatically aggregates and sorts all your coding notes from the `note/` directory by difficulty (**Easy**, **Medium**, and **Hard**).
---
## 📂 Difficulty Categorization
> [!SUCCESS] 🟢 Easy Problems
> List of all Easy difficulty questions.
> ```dataview
> TABLE choice(status = "Solved", "🟢 Solved", "🔴 Unsolved") AS Status
> FROM "leetcode/note/easy" OR #leetcode/easy
> SORT file.name ASC
> ```
> [!WARNING] 🟡 Medium Problems
> List of all Medium difficulty questions.
> ```dataview
> TABLE choice(status = "Solved", "🟢 Solved", "🔴 Unsolved") AS Status
> FROM "leetcode/note/medium" OR #leetcode/medium
> SORT file.name ASC
> ```
> [!DANGER] 🔴 Hard Problems
> List of all Hard difficulty questions.
> ```dataview
> TABLE choice(status = "Solved", "🟢 Solved", "🔴 Unsolved") AS Status
> FROM "leetcode/note/hard" OR #leetcode/hard
> SORT file.name ASC
> ```
@@ -0,0 +1,60 @@
---
id: 1346
title: Check If N and Its Double Exist
difficulty: Easy
tags:
- array
- hash-table
- two-pointers
- binary-search
- sorting
status: Solved
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/check-if-n-and-its-double-exist/
review_needed: false
---
# 1346. Check If N and Its Double Exist
> [!info] **Problem Link**: [LeetCode - Check If N and Its Double Exist](https://leetcode.com/problems/check-if-n-and-its-double-exist/)
## 📝 Problem Description
Given an array `arr` of integers, check if there exist two indices `i` and `j` such that :
- ` i != j`
- `0 <= i, j < arr.length`
- `arr[i] == 2 * arr[j]`
---
### 📥 Example 1
> **Input:** `arr = [10,2,5,3]`
> **Output:** `true`
> **Explanation:** For `i = 0` and `j = 2`, `arr[i] == 10 == 2 * 5 == 2 * arr[j]`
### 📥 Example 2
> **Input:** `arr = [3,1,7,11]`
> **Output:** `false`
> **Explanation:** There is no i and j that satisfy the conditions.
---
## 💡 Approaches & Explanations
Have a mem that hold int that you have see before. for each int in the array you check if you seen double or half in the mem if it is than return True. In the end return False
## 💻 Code Implementations
### Python3
```python
class Solution:
def checkIfExist(self, arr: List[int]) -> bool:
mem = []
for idx , i in enumerate(arr):
if i*2 in mem or i/2 in mem:
return True
mem.append(i)
return False
```
@@ -0,0 +1,58 @@
---
id: 169
title: Majority Element
difficulty: Easy
tags:
- array
- hash-table
- sorting
- counting
status: Solved
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/majority-element/
review_needed: false
---
# 169. Majority Element
> [!info] **Problem Link**: [LeetCode - Majority Element](https://leetcode.com/problems/majority-element/)
## 📝 Problem Description
Given an array `nums` of size `n`, return the majority element.
The majority element is the element that appears more than `[n / 2]` times. You may assume that the majority element always exists in the array.
---
### 📥 Example 1
> **Input:** `nums = [3,2,3]`
> **Output:** `3`
### 📥 Example 2
> **Input:** `nums = [2,2,1,1,1,2,2]`
> **Output:** `2`
---
## 💡 Approaches & Explanations
have cont, if cont is equal to 0 the res become the highest amount
if i equal to the res than add one to count anything else cont -1
return res at the end
## 💻 Code Implementations
### Python3
```python
class Solution:
def majorityElement(self, nums: List[int]) -> int:
cont = 0
res = None
for i in nums:
if cont == 0:
res = i
cont +=1 if res == i else -1
return res
```
+82
View File
@@ -0,0 +1,82 @@
---
id: 1
title: Two Sum
difficulty: Easy
tags:
- array
- hash-table
status: Solved
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/two-sum/
review_needed: false
---
# 1. Two Sum
> [!info] **Problem Link**: [LeetCode - Two Sum](https://leetcode.com/problems/two-sum/)
## 📝 Problem Description
Given an array of integers `nums` and an integer `target`, return *indices of the two numbers such that they add up to `target`*.
You may assume that each input would have ***exactly* one solution**, and you may not use the *same* element twice.
You can return the answer in any order.
---
### 📥 Example 1
> **Input:** `nums = [2,7,11,15]`, `target = 9`
> **Output:** `[0,1]`
> **Explanation:** Because `nums[0] + nums[1] == 9`, we return `[0, 1]`.
### 📥 Example 2
> **Input:** `nums = [3,2,4]`, `target = 6`
> **Output:** `[1,2]`
### 📥 Example 3
> **Input:** `nums = [3,3]`, `target = 6`
> **Output:** `[0,1]`
---
## 💡 Approaches & Explanations
### Approach 1: Hash Map (One-Pass) — *Optimal*
The optimal approach is to use a hash map to keep track of the numbers we have seen so far and their indices. As we iterate through the array, we check if the complement (`target - nums[i]`) already exists in our hash map.
- If it does, we found the pair and return their indices.
- If it doesn't, we add the current number and its index to the hash map.
#### 📊 Complexity Analysis
- **Time Complexity:** $\mathcal{O}(N)$ where $N$ is the number of elements in the array. We traverse the list containing $N$ elements only once, and lookup in the hash table takes $\mathcal{O}(1)$ time.
- **Space Complexity:** $\mathcal{O}(N)$ since we store at most $N$ elements in the hash map.
---
### Approach 2: Brute Force
Compare every pair of numbers to see if their sum equals the target.
- **Time Complexity:** $\mathcal{O}(N^2)$
- **Space Complexity:** $\mathcal{O}(1)$
---
## 💻 Code Implementations
### Python3
```python
class Solution:
def twoSum(self, nums: List[int], target: int) -> List[int]:
seen = {} # val -> index
for i, num in enumerate(nums):
complement = target - num
if complement in seen:
return [seen[complement], i]
seen[num] = i
return []
```
---
## 🧠 Key Takeaways & Lessons
- **The Complement Trick:** When looking for a pair that sums to a target, rephrase the search: instead of looking for $A + B = \text{target}$, look for $\text{complement} = \text{target} - A$ that is already stored.
- **Hash Map for $\mathcal{O}(1)$ Lookups:** Trading memory (space complexity) for time complexity is a common pattern in array search problems.
@@ -0,0 +1,88 @@
---
id: 217
title: Contains Duplicate
difficulty: Easy
tags:
- array
- hash-table
status: Solved
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/contains-duplicate/
review_needed: false
---
# 217. Contains Duplicate
> [!info] **Problem Link**: [LeetCode - Contains Duplicate](https://leetcode.com/problems/contains-duplicate/)
## 📝 Problem Description
Given an integer array `nums`, return `true` if *any value appears at least twice in the array*, and return `false` if every element is distinct.
---
### 📥 Example 1
> **Input:** `nums = [1,2,3,1]`
> **Output:** `true`
> **Explanation:** The element 1 occurs at the indices 0 and 3.
### 📥 Example 2
> **Input:** `nums = [1,2,3,4]`
> **Output:** `false`
### 📥 Example 3
> **Input:** `nums = [1,1,1,3,3,4,3,2,4,2]`
> **Output:** `true`
---
## 💡 Approaches & Explanations
### Approach 1: Hash Set (Length Comparison)
The simplest way to check for duplicates in Python is to convert the array `nums` into a set. A set only contains unique elements, so:
- If there are duplicates, the length of the set will be less than the length of the array.
- If all elements are unique, the lengths will be equal.
#### 📊 Complexity Analysis
- **Time Complexity:** $\mathcal{O}(N)$ where $N$ is the number of elements in the array. Converting an array to a set requires traversing the entire array and inserting each element.
- **Space Complexity:** $\mathcal{O}(N)$ as we store up to $N$ unique elements in the set.
---
### Approach 2: Hash Set (Early Return / One-Pass) — *Alternative*
Instead of converting the entire array to a set, we can iterate through the array and store elements in a set as we go. If we encounter an element that is already in the set, we can return `true` immediately. This avoids processing the rest of the array.
#### 📊 Complexity Analysis
- **Time Complexity:** $\mathcal{O}(N)$ in the worst case (no duplicates). In the best case, it can be $\mathcal{O}(1)$ if a duplicate is found at the beginning.
- **Space Complexity:** $\mathcal{O}(N)$ to store the visited elements.
---
## 💻 Code Implementations
### Python3
#### Option A: Length Comparison (Concise)
```python
class Solution:
def containsDuplicate(self, nums: List[int]) -> bool:
return len(nums) != len(set(nums))
```
#### Option B: Early Return (Optimal for large lists with early duplicates)
```python
class Solution:
def containsDuplicate(self, nums: List[int]) -> bool:
seen = set()
for num in nums:
if num in seen:
return True
seen.add(num)
return False
```
---
## 🧠 Key Takeaways & Lessons
- **Hash Set for Uniqueness:** Sets are the go-to data structure when you need to verify uniqueness or look up elements in $\mathcal{O}(1)$ time.
- **Early Return Optimization:** While converting the whole list to a set is clean and concise, iterating and returning early when a duplicate is found can save time and memory in practice.
- **Time-Space Trade-off:** We use extra space ($\mathcal{O}(N)$ memory) to achieve linear time complexity ($\mathcal{O}(N)$) instead of a brute-force search ($\mathcal{O}(N^2)$).
+13
View File
@@ -0,0 +1,13 @@
---
id: 242
title: Valid Anagram
difficulty: Easy
tags:
- hash-table
- string
- sorting
status: unsolve
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/valid-anagram/
review_needed: false
---
+12
View File
@@ -0,0 +1,12 @@
---
id: 290
title: Word Pattern
difficulty: Easy
tags:
- hash-table
- string
status: unsolve
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/word-pattern/
review_needed: false
---
@@ -0,0 +1,12 @@
---
id: 724
title: Find Pivot Index
difficulty: Easy
tags:
- array
- prefix-sum
status: unsolve
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/find-pivot-index/
review_needed: false
---
@@ -0,0 +1,13 @@
---
id: 30
title: Substring with Concatenation of All Words
difficulty: Hard
tags:
- hash-table
- string
- sliding-window
status: unsolve
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/substring-with-concatenation-of-all-words/
review_needed: false
---
@@ -0,0 +1,15 @@
---
id: 381
title: Insert Delete GetRandom O(1) - Duplicates allowed
difficulty: Hard
tags:
- array
- hash-table
- math
- randomized
- design
status: unsolve
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/insert-delete-getrandom-o1-duplicates-allowed/
review_needed: false
---
@@ -0,0 +1,12 @@
---
id: 41
title: First Missing Positive
difficulty: Hard
tags:
- array
- hash-table
status: unsolve
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/first-missing-positive/
review_needed: false
---
@@ -0,0 +1,13 @@
---
id: 128
title: Longest Consecutive Sequence
difficulty: Medium
tags:
- array
- hash-table
- union-find
status: unsolve
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/longest-consecutive-sequence/
review_needed: false
---
+13
View File
@@ -0,0 +1,13 @@
---
id: 2017
title: Grid Game
difficulty: Medium
tags:
- array
- matrix
- prefix-sum
status: unsolve
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/grid-game/
review_needed: false
---
@@ -0,0 +1,12 @@
---
id: 238
title: Product of Array Except Self
difficulty: Medium
tags:
- array
- prefix-sum
status: unsolve
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/product-of-array-except-self/
review_needed: false
---
@@ -0,0 +1,14 @@
---
id: 271
title: Encode and Decode Strings
difficulty: Medium
tags:
- array
- hash-table
- string
- design
status: unsolve
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/encode-and-decode-strings/
review_needed: false
---
@@ -0,0 +1,17 @@
---
id: 347
title: Top K Frequent Elements
difficulty: Medium
tags:
- array
- hash-table
- divide-and-conquer
- sorting
- heap-priority-queue
- bucket-sort
- quickselect
status: unsolve
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/top-k-frequent-elements/
review_needed: false
---
+13
View File
@@ -0,0 +1,13 @@
---
id: 36
title: Valid Sudoku
difficulty: Medium
tags:
- array
- hash-table
- matrix
status: unsolve
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/valid-sudoku/
review_needed: false
---
@@ -0,0 +1,15 @@
---
id: 380
title: Insert Delete GetRandom O(1)
difficulty: Medium
tags:
- array
- hash-table
- math
- randomized
- design
status: unsolve
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/insert-delete-getrandom-o1/
review_needed: false
---
@@ -0,0 +1,12 @@
---
id: 442
title: Find All Duplicates in an Array
difficulty: Medium
tags:
- array
- hash-table
status: unsolve
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/find-all-duplicates-in-an-array/
review_needed: false
---
+14
View File
@@ -0,0 +1,14 @@
---
id: 49
title: Group Anagrams
difficulty: Medium
tags:
- array
- hash-table
- string
- sorting
status: unsolve
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/group-anagrams/
review_needed: false
---
@@ -0,0 +1,13 @@
---
id: 560
title: Subarray Sum Equals K
difficulty: Medium
tags:
- array
- hash-table
- prefix-sum
status: unsolve
date_solved: 2026-05-26
leetcode_url: https://leetcode.com/problems/subarray-sum-equals-k/
review_needed: false
---
@@ -0,0 +1,375 @@
---
tags:
- leetcode/array-hashing
- study-guide
- coding-practice
difficulty_distribution:
easy: 7
medium: 10
hard: 3
total_problems: 20
completed_problems: 1
progress_percentage: 5%
last_updated: 2026-05-26
---
# 🚀 Array & Hashing: LeetCode Master List
Welcome to your study guide for **Array and Hashing**! This topic is the bedrock of coding interviews, establishing the core patterns for lookups, frequency counting, prefix arrays, and space-time tradeoffs.
---
## 📊 Progress Tracker
| Status | # | Problem | Difficulty | Key Technique | Links |
| :----: | :--: | :------------------------------------------------------------------------------------------------------------------------ | :--------: | :----------------------------- | :--------------------------------------------------------------------------------------: |
| [x] | 1 | [Two Sum](https://leetcode.com/problems/two-sum/) | 🟢 Easy | Hash Map Complement | [LeetCode](https://leetcode.com/problems/two-sum/) |
| [ ] | 217 | [Contains Duplicate](https://leetcode.com/problems/contains-duplicate/) | 🟢 Easy | Hash Set Presence | [LeetCode](https://leetcode.com/problems/contains-duplicate/) |
| [ ] | 242 | [Valid Anagram](https://leetcode.com/problems/valid-anagram/) | 🟢 Easy | Frequency Count / Sorting | [LeetCode](https://leetcode.com/problems/valid-anagram/) |
| [ ] | 169 | [Majority Element](https://leetcode.com/problems/majority-element/) | 🟢 Easy | Boyer-Moore Voting / Hash Map | [LeetCode](https://leetcode.com/problems/majority-element/) |
| [ ] | 290 | [Word Pattern](https://leetcode.com/problems/word-pattern/) | 🟢 Easy | Bijective Hash Mapping | [LeetCode](https://leetcode.com/problems/word-pattern/) |
| [ ] | 1346 | [Check If N and Its Double Exist](https://leetcode.com/problems/check-if-n-and-its-double-exist/) | 🟢 Easy | Hash Set Double-Lookup | [LeetCode](https://leetcode.com/problems/check-if-n-and-its-double-exist/) |
| [ ] | 724 | [Find Pivot Index](https://leetcode.com/problems/find-pivot-index/) | 🟢 Easy | Prefix Sum Balance | [LeetCode](https://leetcode.com/problems/find-pivot-index/) |
| [ ] | 49 | [Group Anagrams](https://leetcode.com/problems/group-anagrams/) | 🟡 Medium | Categorization Key Mapping | [LeetCode](https://leetcode.com/problems/group-anagrams/) |
| [ ] | 347 | [Top K Frequent Elements](https://leetcode.com/problems/top-k-frequent-elements/) | 🟡 Medium | Bucket Sort / Heap Tracking | [LeetCode](https://leetcode.com/problems/top-k-frequent-elements/) |
| [ ] | 238 | [Product of Array Except Self](https://leetcode.com/problems/product-of-array-except-self/) | 🟡 Medium | Left/Right Prefix Products | [LeetCode](https://leetcode.com/problems/product-of-array-except-self/) |
| [ ] | 36 | [Valid Sudoku](https://leetcode.com/problems/valid-sudoku/) | 🟡 Medium | Sub-grid Hashing (Bitmask/Set) | [LeetCode](https://leetcode.com/problems/valid-sudoku/) |
| [ ] | 128 | [Longest Consecutive Sequence](https://leetcode.com/problems/longest-consecutive-sequence/) | 🟡 Medium | Hash Set Boundary Scan | [LeetCode](https://leetcode.com/problems/longest-consecutive-sequence/) |
| [ ] | 560 | [Subarray Sum Equals K](https://leetcode.com/problems/subarray-sum-equals-k/) | 🟡 Medium | Prefix Sum + Frequency Map | [LeetCode](https://leetcode.com/problems/subarray-sum-equals-k/) |
| [ ] | 271 | [Encode and Decode Strings](https://leetcode.com/problems/encode-and-decode-strings/) | 🟡 Medium | Length-Prefix Chunking | [LeetCode](https://leetcode.com/problems/encode-and-decode-strings/) |
| [ ] | 442 | [Find All Duplicates in an Array](https://leetcode.com/problems/find-all-duplicates-in-an-array/) | 🟡 Medium | In-place Sign Negation | [LeetCode](https://leetcode.com/problems/find-all-duplicates-in-an-array/) |
| [ ] | 380 | [Insert Delete GetRandom O(1)](https://leetcode.com/problems/insert-delete-getrandom-o1/) | 🟡 Medium | Array + Map Index Swapping | [LeetCode](https://leetcode.com/problems/insert-delete-getrandom-o1/) |
| [ ] | 2017 | [Grid Game](https://leetcode.com/problems/grid-game/) | 🟡 Medium | 2-Row Prefix Sum Selection | [LeetCode](https://leetcode.com/problems/grid-game/) |
| [ ] | 41 | [First Missing Positive](https://leetcode.com/problems/first-missing-positive/) | 🔴 Hard | Cyclic In-place Sorting | [LeetCode](https://leetcode.com/problems/first-missing-positive/) |
| [ ] | 30 | [Substring with Concatenation](https://leetcode.com/problems/substring-with-concatenation-of-all-words/) | 🔴 Hard | Sliding Window + Count Maps | [LeetCode](https://leetcode.com/problems/substring-with-concatenation-of-all-words/) |
| [ ] | 381 | [Insert Delete GetRandom O(1) - Duplicates](https://leetcode.com/problems/insert-delete-getrandom-o1-duplicates-allowed/) | 🔴 Hard | Array + Map to Set of Indices | [LeetCode](https://leetcode.com/problems/insert-delete-getrandom-o1-duplicates-allowed/) |
---
## 🟢 Easy Problems (7)
### 1. Two Sum (LC 1)
> [!TIP]
> **Core Concept:** Instead of checking all pairs, store the numbers you have seen in a hash map mapping `value -> index`. For each `num`, check if its complement (`target - num`) is already in the map.
- **Complexity Target:** $O(n)$ Time | $O(n)$ Space
- **Starter Template:**
```python
def twoSum(nums: list[int], target: int) -> list[int]:
pass
```
- **My Notes / Solution:**
*
---
### 2. Contains Duplicate (LC 217)
> [!TIP]
> **Core Concept:** Iterate through the array and store each element in a Hash Set. If the element is already in the set, a duplicate has been found.
- **Complexity Target:** $O(n)$ Time | $O(n)$ Space
- **Starter Template:**
```python
def containsDuplicate(nums: list[int]) -> bool:
pass
```
- **My Notes / Solution:**
*
---
### 3. Valid Anagram (LC 242)
> [!TIP]
> **Core Concept:** Count the frequency of characters in both strings. You can use two hash maps, or a single array size 26 if constraints are purely lowercase a-z. Compare the frequency distributions.
- **Complexity Target:** $O(n)$ Time | $O(1)$ Space (since lowercase English alphabet count is fixed at 26)
- **Starter Template:**
```python
def isAnagram(s: str, t: str) -> bool:
pass
```
- **My Notes / Solution:**
*
---
### 4. Majority Element (LC 169)
> [!TIP]
> **Core Concept:** While a Hash Map works, you can solve this in $O(1)$ space using **Boyer-Moore Voting Algorithm**. Keep a `candidate` and a `count`. When `count == 0`, pick the current element as candidate. Increment `count` if the element matches the candidate, else decrement it.
- **Complexity Target:** $O(n)$ Time | $O(1)$ Space
- **Starter Template:**
```python
def majorityElement(nums: list[int]) -> int:
pass
```
- **My Notes / Solution:**
*
---
### 5. Word Pattern (LC 290)
> [!TIP]
> **Core Concept:** Establish a bijective (two-way) mapping between character in `pattern` and words in string `s` using two hash maps. If a key maps to a different value in either map, return `False`.
- **Complexity Target:** $O(n)$ Time | $O(n)$ Space (where $n$ is total chars/words)
- **Starter Template:**
```python
def wordPattern(pattern: str, s: str) -> bool:
pass
```
- **My Notes / Solution:**
*
---
### 6. Check If N and Its Double Exist (LC 1346)
> [!TIP]
> **Core Concept:** Iterate through the array. Check if `2 * num` or `num / 2` (if divisible by 2) exists in your Hash Set of previously visited values.
- **Complexity Target:** $O(n)$ Time | $O(n)$ Space
- **Starter Template:**
```python
def checkIfExist(arr: list[int]) -> bool:
pass
```
- **My Notes / Solution:**
*
---
### 7. Find Pivot Index (LC 724)
> [!TIP]
> **Core Concept:** Compute the total sum of the array. Track the running `left_sum`. For each index, check if `left_sum == total_sum - left_sum - num`. If it is, that's the pivot.
- **Complexity Target:** $O(n)$ Time | $O(1)$ Space
- **Starter Template:**
```python
def pivotIndex(nums: list[int]) -> int:
pass
```
- **My Notes / Solution:**
*
---
## 🟡 Medium Problems (10)
### 8. Group Anagrams (LC 49)
> [!TIP]
> **Core Concept:** Group words by their character signature. The signature can be a sorted string, or a character count tuple `[0] * 26` mapped to list of matching strings.
- **Complexity Target:** $O(n \cdot m)$ Time (where $m$ is max word length) | $O(n \cdot m)$ Space
- **Starter Template:**
```python
def groupAnagrams(strs: list[str]) -> list[list[str]]:
pass
```
- **My Notes / Solution:**
*
---
### 9. Top K Frequent Elements (LC 347)
> [!TIP]
> **Core Concept:** Count frequencies using a hash map. Instead of sorting (which is $O(n \log n)$), use **Bucket Sort** where index represents frequencies. Since the max frequency is capped at `len(nums)`, we can assemble the result in linear time.
- **Complexity Target:** $O(n)$ Time | $O(n)$ Space
- **Starter Template:**
```python
def topKFrequent(nums: list[int], k: int) -> list[int]:
pass
```
- **My Notes / Solution:**
*
---
### 10. Product of Array Except Self (LC 238)
> [!TIP]
> **Core Concept:** Create an output array. Do a forward pass to store the prefix product at each index. Then do a backward pass, keeping a running suffix product, multiplying it into your result.
- **Complexity Target:** $O(n)$ Time | $O(1)$ Extra Space (excluding output array)
- **Starter Template:**
```python
def productExceptSelf(nums: list[int]) -> list[int]:
pass
```
- **My Notes / Solution:**
*
---
### 11. Valid Sudoku (LC 36)
> [!TIP]
> **Core Concept:** Track duplicates in rows, columns, and 3x3 sub-grids. Use a set for each row, column, and sub-grid. The subgrid index can be identified using `(r // 3, c // 3)`.
- **Complexity Target:** $O(1)$ Time & Space (since grid size is constant $9 \times 9$)
- **Starter Template:**
```python
def isValidSudoku(board: list[list[str]]) -> bool:
pass
```
- **My Notes / Solution:**
*
---
### 12. Longest Consecutive Sequence (LC 128)
> [!TIP]
> **Core Concept:** Convert the array to a Hash Set. Loop through each number; if `num - 1` is not in the set, it means `num` is the *start* of a new sequence. Scan forward (`num + 1`, `num + 2`...) to find the sequence length. This guarantees each element is processed at most twice.
- **Complexity Target:** $O(n)$ Time | $O(n)$ Space
- **Starter Template:**
```python
def longestConsecutive(nums: list[int]) -> int:
pass
```
- **My Notes / Solution:**
*
---
### 13. Subarray Sum Equals K (LC 560)
> [!TIP]
> **Core Concept:** Use prefix sum properties. If the difference between current prefix sum and target `k` (i.e. `prefix_sum - k`) was seen previously as a prefix sum, the subarray between those indices sums to `k`. Store prefix sum frequencies in a map.
- **Complexity Target:** $O(n)$ Time | $O(n)$ Space
- **Starter Template:**
```python
def subarraySum(nums: list[int], k: int) -> int:
pass
```
- **My Notes / Solution:**
*
---
### 14. Encode and Decode Strings (LC 271 / Premium)
> [!TIP]
> **Core Concept:** To combine strings safely, prepend each string with its length followed by a delimiter (e.g. `"4#neet"`). When decoding, parse the number, jump past the delimiter, and slice that exact length.
- **Complexity Target:** $O(n)$ Time | $O(1)$ Extra Space (excluding output list)
- **Starter Template:**
```python
class Codec:
def encode(self, strs: list[str]) -> str:
pass
def decode(self, s: str) -> list[str]:
pass
```
- **My Notes / Solution:**
*
---
### 15. Find All Duplicates in an Array (LC 442)
> [!TIP]
> **Core Concept:** The array contains elements from `1` to `n`. You can use the values as indices. As you iterate, look up `nums[abs(x) - 1]`. If it is positive, negate it. If it is already negative, then `abs(x)` has been seen before.
- **Complexity Target:** $O(n)$ Time | $O(1)$ Extra Space
- **Starter Template:**
```python
def findDuplicates(nums: list[int]) -> list[int]:
pass
```
- **My Notes / Solution:**
*
---
### 16. Insert Delete GetRandom O(1) (LC 380)
> [!TIP]
> **Core Concept:** Store values in an Array List to achieve $O(1)$ random lookups. Maintain a Hash Map `val -> index` to achieve $O(1)$ insertions and updates. When deleting, swap the element to delete with the last element in the array, update the map, and pop from the array.
- **Complexity Target:** $O(1)$ Time average | $O(n)$ Space
- **Starter Template:**
```python
class RandomizedSet:
def __init__(self):
pass
def insert(self, val: int) -> bool:
pass
def remove(self, val: int) -> bool:
pass
def getRandom(self) -> int:
pass
```
- **My Notes / Solution:**
*
---
### 17. Grid Game (LC 2017)
> [!TIP]
> **Core Concept:** The first robot splits the grid into two paths. Compute prefix sums for row 0 and row 1. Robot 1 wants to minimize Robot 2's maximum score, which will always be the max of the remaining top-right elements or bottom-left elements after Robot 1 pivots.
- **Complexity Target:** $O(n)$ Time | $O(n)$ Space (or $O(1)$ if prefix sums are computed on-the-fly)
- **Starter Template:**
```python
def gridGame(grid: list[list[int]]) -> int:
pass
```
- **My Notes / Solution:**
*
---
## 🔴 Hard Problems (3)
### 18. First Missing Positive (LC 41)
> [!TIP]
> **Core Concept:** Use **Cyclic Sort** style placement. Iterate through the array and try to place each number `x` (if it lies in range `[1, n]`) at its correct index `x - 1`. Perform this swap in a loop. Finally, scan the array to find the first index `i` where `nums[i] != i + 1`.
- **Complexity Target:** $O(n)$ Time | $O(1)$ Space
- **Starter Template:**
```python
def firstMissingPositive(nums: list[int]) -> int:
pass
```
- **My Notes / Solution:**
*
---
### 19. Substring with Concatenation of All Words (LC 30)
> [!TIP]
> **Core Concept:** All words are of equal length `L`. Use a sliding window starting at each offset `0 <= i < L`. For each window, use a hash map to count occurrences of words of length `L` and compare it with the count map of the target `words` list.
- **Complexity Target:** $O(n \cdot L)$ Time (where $n$ is string length, $L$ is word length) | $O(m \cdot L)$ Space (where $m$ is number of words)
- **Starter Template:**
```python
def findSubstring(s: str, words: list[str]) -> list[int]:
pass
```
- **My Notes / Solution:**
*
---
### 20. Insert Delete GetRandom O(1) - Duplicates Allowed (LC 381)
> [!TIP]
> **Core Concept:** Similar to LC 380, but the Hash Map now maps `val -> Set of indices` where the value resides in our dynamic list. During deletion, lookup any index from the value's set, swap with the list's last element, and update the set of the swapped element accordingly.
- **Complexity Target:** $O(1)$ Time average | $O(n)$ Space
- **Starter Template:**
```python
class RandomizedCollection:
def __init__(self):
pass
def insert(self, val: int) -> bool:
pass
def remove(self, val: int) -> bool:
pass
def getRandom(self) -> int:
pass
```
- **My Notes / Solution:**
*
+237
View File
@@ -0,0 +1,237 @@
---
title: Introduction to NumPy
tags:
- python
- numpy
- data-science
- numerical-computing
category: Lesson
difficulty: Beginner to Intermediate
created: 2026-05-26
---
# Introduction to NumPy
> [!abstract] What is NumPy?
> **NumPy** (Numerical Python) is the foundational library for scientific computing in Python. It provides a high-performance multidimensional array object, tools for working with these arrays, and linear algebra, Fourier transform, and random number capabilities.
>
> ### Why use NumPy instead of standard Python lists?
> 1. **Speed:** NumPy arrays are written in C, making mathematical operations up to 100x faster than standard Python lists.
> 2. **Memory Efficiency:** NumPy arrays use contiguous blocks of memory, whereas Python lists store pointers to objects scattered across memory.
> 3. **Vectorization:** It allows performing mathematical operations on whole arrays without writing slow `for` loops.
---
## 📐 The N-Dimensional Array (`ndarray`)
The core of NumPy is the **`ndarray`** (N-dimensional array). It is a grid of values, all of the **same type** (homogenous), indexed by a tuple of non-negative integers.
```mermaid
graph TD
A[ndarray] --> B["1D Array (Vector) <br> Shape: (n,)"]
A --> C["2D Array (Matrix) <br> Shape: (m, n)"]
A --> D["3D Array (Tensor) <br> Shape: (p, m, n)"]
```
### Essential Array Attributes
Every array has attributes that describe its structure:
```python
import numpy as np
arr = np.array([[1, 2, 3], [4, 5, 6]])
print(arr.ndim) # Number of dimensions (axes) -> 2
print(arr.shape) # Tuple representing sizes in each dimension -> (2, 3)
print(arr.size) # Total number of elements -> 6
print(arr.dtype) # Data type of the elements -> int64
```
---
## 🛠️ Creating Arrays
First, import the library using the standard alias:
```python
import numpy as np
```
### 1. From Python Lists
```python
# 1D Vector
v = np.array([1, 2, 3])
# 2D Matrix
m = np.array([[1, 2], [3, 4]])
```
### 2. Built-in Placeholders
NumPy provides functions to initialize arrays with placeholders, avoiding manual creation:
```python
# Array of zeros
zeros = np.zeros((3, 4)) # 3 rows, 4 columns
# Array of ones
ones = np.ones((2, 3), dtype=np.int32)
# Range of numbers (similar to range())
range_arr = np.arange(0, 10, 2) # [0, 2, 4, 6, 8]
# Linearly spaced numbers
linspace_arr = np.linspace(0, 1, 5) # [0.0, 0.25, 0.5, 0.75, 1.0]
# Identity Matrix
eye_matrix = np.eye(3) # 3x3 identity matrix
```
### 3. Random Number Generation
```python
# Uniform random values between [0.0, 1.0)
rand_arr = np.random.rand(2, 2)
# Standard normal distribution (mean=0, std=1)
randn_arr = np.random.randn(2, 2)
# Random integers
rand_ints = np.random.randint(1, 100, size=(5,))
```
---
## ⚡ Element-wise Operations & Vectorization
In standard Python, to add two lists element-wise, you need a list comprehension or loop. In NumPy, you do it directly.
```python
x = np.array([1, 2, 3])
y = np.array([4, 5, 6])
print(x + y) # [5, 7, 9]
print(x * y) # [4, 10, 18]
print(x ** 2) # [1, 4, 9]
```
### 📡 Broadcasting
Broadcasting is a powerful mechanism that allows NumPy to perform arithmetic operations on arrays of **different shapes**. The smaller array is "broadcast" across the larger array so that they have compatible shapes.
```python
matrix = np.array([[1, 2, 3], [4, 5, 6]])
scalar = 10
# The scalar is added to every single element
print(matrix + scalar)
# [[11, 12, 13]
# [14, 15, 16]]
```
---
## 🔍 Indexing, Slicing & Masking
### Slicing 2D Arrays
Slicing follows the format `array[row_start:row_end, col_start:col_end]`.
```python
arr = np.array([
[10, 11, 12],
[20, 21, 22],
[30, 31, 32]
])
# Get row at index 1
print(arr[1, :]) # [20, 21, 22]
# Get column at index 2
print(arr[:, 2]) # [12, 22, 32]
# Slice a subgrid (top-left 2x2)
print(arr[0:2, 0:2])
# [[10, 11]
# [20, 21]]
```
### 🎭 Boolean Masking (Conditional Filtering)
You can filter arrays using conditions. NumPy returns elements where the condition resolves to `True`.
```python
data = np.array([1, 5, 8, 12, 3, 15])
# Create a boolean mask
mask = data > 5 # [False, False, True, True, False, True]
# Filter using the mask
filtered_data = data[mask] # [8, 12, 15]
```
---
## 🧮 Common Aggregations & Axis Operations
Aggregations allow you to compute statistics over entire arrays or along specific **axes**:
* `axis=0`: Down the columns (collapses rows).
* `axis=1`: Across the rows (collapses columns).
```python
arr = np.array([[1, 2], [3, 4]])
# Sum of all elements
print(np.sum(arr)) # 10
# Sum down the columns (vertical)
print(np.sum(arr, axis=0)) # [4, 6]
# Sum across the rows (horizontal)
print(np.sum(arr, axis=1)) # [3, 7]
```
---
## 🔄 Reshaping and Transposing
You can change the shape of an array without changing its data using `.reshape()` or `.T` (Transpose).
```python
flat = np.arange(1, 7) # [1, 2, 3, 4, 5, 6]
# Reshape into a 2x3 matrix
matrix = flat.reshape(2, 3)
# [[1, 2, 3]
# [4, 5, 6]]
# Transpose matrix (swap rows and columns)
transposed = matrix.T
# [[1, 4]
# [2, 5]
# [3, 6]]
```
---
## 💡 Best Practices
> [!important] Avoid Standard Loops
> Standard loops in Python are interpreted, which adds massive overhead. Vectorized operations execute in compiled C, taking advantage of CPU caches and SIMD instructions.
>
> **Example Comparison:**
> ```python
> # ❌ Extremely Slow
> values = np.random.rand(1_000_000)
> reciprocal = [1 / x for x in values]
>
> # ✅ Near Instantaneous
> reciprocal = 1 / values
> ```
> [!tip] Use In-place Operations to Save Memory
> Instead of creating a new copy, perform calculations directly on the existing array if possible using syntax like `+=`, `-=`, or `*=`.
> ```python
> a = np.ones(1000000)
> b = np.ones(1000000)
>
> # Allocates new memory
> a = a + b
>
> # Modifies 'a' in-place (saves memory allocation time)
> a += b
> ```
+241
View File
@@ -0,0 +1,241 @@
---
title: Introduction to Pandas
tags:
- python
- pandas
- data-science
- data-analysis
category: Lesson
difficulty: Beginner to Intermediate
created: 2026-05-26
---
# Introduction to Pandas
> [!abstract] What is Pandas?
> **Pandas** is a fast, powerful, flexible, and easy-to-use open-source data analysis and manipulation library built on top of the Python programming language.
>
> The name is derived from **"Panel Data"**, an econometrics term for multidimensional structured data sets. It is the foundational library for Data Science, Machine Learning, and Data Analysis in Python, acting as the bridge between raw data files (like CSVs, Excel files, or SQL databases) and numerical/modeling libraries (like NumPy, Scikit-Learn, and PyTorch).
---
## 🏗️ Core Data Structures
Pandas simplifies data manipulation by providing two primary, highly optimized data structures:
```mermaid
graph TD
A[Pandas Data Structures] --> B[Series - 1D]
A --> C[DataFrame - 2D]
B -->|Multiple columns merged| C
C -->|Single column extracted| B
```
### 1. Series (1D)
A **Series** is a one-dimensional array-like object containing an array of data and an associated array of data labels, called its **index**. Think of it as a single column in a spreadsheet.
```python
import pandas as pd
# Creating a Series
temperatures = pd.Series([22.5, 24.0, 19.5, 21.8], name="Temp")
print(temperatures)
```
### 2. DataFrame (2D)
A **DataFrame** represents a tabular, spreadsheet-like data structure containing an ordered collection of columns, each of which can be a different value type (numeric, string, boolean, etc.). It has both a row index and a column index.
| Index | Name | Age | Department |
| :--- | :--- | :--- | :--- |
| **0** | Alice | 28 | Engineering |
| **1** | Bob | 34 | Marketing |
| **2** | Charlie | 22 | HR |
---
## 🛠️ Getting Started & Creating Data
First, make sure you have pandas imported. The universal convention is to alias it as `pd`.
```python
import pandas as pd
import numpy as np
```
### Creating DataFrames Manually
You can easily create DataFrames from dictionaries or lists:
```python
data = {
'Name': ['Alice', 'Bob', 'Charlie', 'David'],
'Age': [25, 30, 35, 40],
'Salary': [70000, 80000, 120000, 90000],
'Department': ['HR', 'Engineering', 'Engineering', 'Finance']
}
df = pd.DataFrame(data)
```
---
## 🔍 Essential Operations (The "Big Five")
### 1. Reading & Writing Data
Pandas supports a wide range of file formats out of the box.
```python
# Reading data
df_csv = pd.read_csv('employees.csv')
df_excel = pd.read_excel('sales.xlsx', sheet_name='Sheet1')
df_sql = pd.read_sql('SELECT * FROM users', database_connection)
# Writing data
df.to_csv('output.csv', index=False) # index=False avoids writing the row numbers
df.to_json('output.json')
```
### 2. Inspecting Your Data
Before doing any analysis, you must understand the shape and types of your data.
```python
df.head(2) # Returns the first 2 rows
df.tail(2) # Returns the last 2 rows
df.info() # Summary of columns, non-null counts, and data types
df.describe() # Generates descriptive statistics for numerical columns
df.shape # Returns (rows, columns) as a tuple
```
### 3. Selection & Filtering
Selecting data is one of the most common tasks. Pandas provides multiple intuitive ways to do this.
#### Column Selection
```python
# Select a single column (returns a Series)
ages = df['Age']
# Select multiple columns (returns a DataFrame)
subset = df[['Name', 'Salary']]
```
#### Row Selection using `.loc` and `.iloc`
* `.loc` is **label-based**: references rows/columns by their row labels or column names.
* `.iloc` is **integer-position-based**: references rows/columns by their 0-indexed positions.
```python
# Select the first row by position
first_row = df.iloc[0]
# Select a cell by row label and column name
val = df.loc[2, 'Name'] # Charlie
```
#### Boolean Indexing (Filtering)
To filter rows based on conditions:
```python
# Filter employees earning more than 85,000
high_earners = df[df['Salary'] > 85000]
# Combine multiple conditions using & (AND) or | (OR)
# Always wrap conditions in parentheses!
eng_seniors = df[(df['Department'] == 'Engineering') & (df['Age'] > 30)]
```
### 4. Data Cleaning
Raw data is rarely perfect. Pandas excels at handling missing values and data type conversions.
```python
# Check for missing values
df.isna().sum()
# Drop rows with missing values
df_clean = df.dropna()
# Fill missing values with a default/placeholder
df['Salary'] = df['Salary'].fillna(df['Salary'].mean())
# Rename columns
df = df.rename(columns={'Name': 'Full Name', 'Age': 'Years'})
```
### 5. Grouping & Aggregation
To summarize data by categories, use the Split-Apply-Combine workflow via `groupby`.
```python
# Calculate the average salary by department
dept_salaries = df.groupby('Department')['Salary'].mean()
# Perform multiple aggregations at once
summary = df.groupby('Department').agg({
'Salary': ['mean', 'min', 'max'],
'Age': 'mean'
})
```
---
## 🎓 Hands-on Practice Walkthrough
Let's walk through a mini-scenario. Imagine we have the following DataFrame of store transactions:
```python
transactions = pd.DataFrame({
'TransactionID': [101, 102, 103, 104, 105],
'Store': ['North', 'South', 'North', 'West', 'South'],
'Amount': [250.50, 150.00, np.nan, 300.25, 450.00],
'ItemCount': [3, 2, 1, 5, 4]
})
```
Let's answer three questions:
### Question 1: Fill missing amounts with the median transaction amount.
```python
median_amount = transactions['Amount'].median() # 275.375
transactions['Amount'] = transactions['Amount'].fillna(median_amount)
```
### Question 2: Find transactions with a total amount greater than 200.
```python
large_tx = transactions[transactions['Amount'] > 200]
```
### Question 3: Find the total revenue (sum of amounts) generated by each store.
```python
store_revenue = transactions.groupby('Store')['Amount'].sum().reset_index()
print(store_revenue)
# Store Amount
# 0 North 525.875
# 1 South 600.00
# 2 West 300.25
```
---
## ⚡ Performance Tips & Best Practices
> [!warning] Don't Loop Over Rows!
> Never write `for index, row in df.iterrows():` unless absolutely necessary. Iteration is extremely slow because it disables pandas' vectorized backend.
>
> **Instead, use Vectorization:**
> ```python
> # ❌ Slow & non-pythonic
> for i in range(len(df)):
> df.loc[i, 'Tax'] = df.loc[i, 'Salary'] * 0.1
>
> # ✅ Fast & vectorised
> df['Tax'] = df['Salary'] * 0.1
> ```
> [!tip] Avoid the SettingWithCopyWarning
> When you slice a DataFrame and then modify it, pandas warns you that you might be editing a temporary copy rather than the original source.
> To prevent this, use `.copy()` when creating a subset you intend to modify:
> ```python
> # ❌ Might trigger SettingWithCopyWarning
> engineering = df[df['Department'] == 'Engineering']
> engineering['Bonus'] = 1000
>
> # ✅ Clean and safe
> engineering = df[df['Department'] == 'Engineering'].copy()
> engineering['Bonus'] = 1000
> ```
@@ -0,0 +1,130 @@
---
title: Learn Machine Learning Like a GENIUS and Not Waste Time
channel: InfiniteCodes
url: https://youtu.be/i_LwzRVP7bg
publish_date: 2024-11-14
tags:
- machine-learning
- study-methodology
- learning-strategies
- productivity
category: Video Summary
rating: 5/5
---
# Learn Machine Learning Like a GENIUS and Not Waste Time
> [!abstract] Executive Summary
> This note summarizes the video guide **"Learn Machine Learning Like a GENIUS and Not Waste Time"** by *InfiniteCodes*. The core message is that mastering machine learning (ML) isn't about memorizing complex models or chasing every new trend. Instead, it relies on building a deep intuition of the fundamentals, practicing active learning through project building, reading official documentation, and avoiding the traps of "tutorial hell" and "vibe coding" (blindly relying on LLMs).
---
## 🗺️ The ML Learning Roadmap (Order of Operations)
Beginners often rush into complex deep learning models before understanding basic concepts. The video advocates for a strict, logical **Order of Operations**:
```mermaid
graph TD
A[1. Mathematics Foundations] -->|Linear Algebra & Calculus| B[2. Exploratory Data Analysis]
B -->|Clean, Wrangling, Feature Eng| C[3. Simple Models First]
C -->|Linear/Logistic Reg, Trees| D[4. Advanced Architectures]
D -->|CNNs, RNNs, Transformers| E[5. Real-world Deployment]
```
### 1. Mathematics Foundations
You don't need a math PhD, but you must build a working intuition of:
* **Linear Algebra**: Understanding how data is represented and manipulated as matrices and vectors.
* **Calculus**: Grasping **Gradient Descent** (derivatives, partial derivatives) as the optimization engine that allows models to learn.
### 2. Exploratory Data Analysis (EDA)
Before fitting any model, you must "interview" your data:
* Use libraries like `pandas` and `numpy` to clean and structure data.
* Perform feature engineering and visualize distributions to identify patterns.
### 3. Simple Models First
Start with highly interpretable, foundational algorithms:
* **Linear Regression** & **Logistic Regression**
* **Decision Trees**
* *Why?* They are faster to train, easier to debug, and provide a baseline for more complex models.
### 4. Advanced Architectures
Only dive into deep learning (CNNs, RNNs, Transformers, etc.) when your specific project requires it and your foundations are solid.
---
## 🧠 The "Genius" Implementation Strategy
To truly master an algorithm, do not just read about it. Use this three-step implementation loop:
```mermaid
flowchart LR
Step1[1. Code from Scratch] --> Step2[2. Use Library]
Step2 --> Step3[3. Apply to Real Data]
```
1. **Code from Scratch**: Implement the core algorithm in raw Python (using only `numpy`) to understand the underlying mathematics.
2. **Use a Library**: Implement the same algorithm using `scikit-learn` to see how it is optimized and structured in production-ready libraries.
3. **Apply to Real Data**: Train both implementations on a dataset you gathered or prepared yourself (avoiding clean "toy" datasets).
> [!example] From Scratch vs. Library Example (Linear Regression)
> Below is a comparison of how you should study an algorithm.
>
> === "From Scratch (Math Intuition)"
> ```python
> import numpy as np
>
> class SimpleLinearRegression:
> def __init__(self, lr=0.01, epochs=1000):
> self.lr = lr
> self.epochs = epochs
> self.weights = None
> self.bias = None
>
> def fit(self, X, y):
> n_samples, n_features = X.shape
> self.weights = np.zeros(n_features)
> self.bias = 0
>
> # Gradient Descent loop
> for _ in range(self.epochs):
> y_predicted = np.dot(X, self.weights) + self.bias
> dw = (1 / n_samples) * np.dot(X.T, (y_predicted - y))
> db = (1 / n_samples) * np.sum(y_predicted - y)
>
> self.weights -= self.lr * dw
> self.bias -= self.lr * db
> ```
>
> === "Using a Library (Production Standard)"
> ```python
> from sklearn.linear_model import LinearRegression
>
> # Initialize and fit
> model = LinearRegression()
> model.fit(X_train, y_train)
>
> # Make predictions
> predictions = model.predict(X_test)
> ```
---
## 🚫 Key Pitfalls to Avoid
> [!danger] 1. The "3-Month Fallacy"
> Avoid courses or tutorials promising machine learning mastery in 3 months. Transitioning to a professional level takes sustained, long-term effort and continuous learning.
> [!warning] 2. Tutorial Hell & "Vibe Coding"
> * **Tutorial Hell**: Mindlessly consuming tutorials without writing code. Rule of thumb: Watch **maximum 2 tutorials** on a topic, then immediately build something yourself.
> * **Vibe Coding**: Relying entirely on LLMs (like Cursor or Copilot) to generate code for you. If you don't understand the lines of code being generated, you are building a house of cards.
> [!important] 3. Documentation First
> Build a habit of reading the **official documentation** (e.g., PyTorch, scikit-learn, Pandas) instead of asking an AI for code snippets immediately. This builds strong neural connections and developer independence.
---
## ⚡ Soft Skills & Mindset
* **Deep Work**: Dedicate uninterrupted **90 to 120-minute** blocks to study and code. Turn off notifications and focus deeply.
* **Problem Solving**: ML is about breaking complex, abstract real-world problems down into structured data steps.
* **Community & Networking**: Share your learning journey on GitHub, LinkedIn, or community Discords. Learning in public accelerates growth and opens career opportunities.
+20
View File
@@ -0,0 +1,20 @@
---
title: LLM Notes Index
created: 2026-05-26
tags:
- llm
- index
- moc
category: index
---
# LLM Notes Index
All notes under `note/LLM/`, grouped by subfolder.
```dataview
LIST rows.file.link
FROM "note/LLM"
WHERE file.name != "index"
GROUP BY file.folder
SORT file.folder ASC
```
+48
View File
@@ -0,0 +1,48 @@
---
title: Learn Linux - The Full Course
channel: Boot dev
url: https://youtu.be/v392lEyM29A?si=2QnHdGS-DdjtS5Bg
tags:
- linux
- course
- devops
- terminal
date: 2026-05-29
---
# Learn Linux - The Full Course
## Overview
This comprehensive course, led by Lane from **Boot.dev**, is designed to provide developers, DevOps engineers, and IT professionals with a solid foundation in Linux and Unix-like systems. The course focuses on moving beyond basic usage to mastering the command line, navigating filesystems, and using powerful CLI tools.
## Key Topics
### Terminals and Shells (04:03)
- Deep dive into the fundamental interface of Linux.
- Explaining the differences between terminals and shells.
- How they facilitate system interaction.
### Filesystems (20:43)
- Navigating the Linux directory structure.
- Managing files and understanding data organization on disk.
### Permissions (51:18)
- Critical look at the Linux security model.
- User and group management.
- Reading and modifying file permissions.
### Programs (01:13:31)
- How software is executed in a Linux environment.
- Process management.
- Configuration of the system `PATH`.
### Input/Output (01:38:33)
- Mastery of data manipulation using standard input/output.
- Redirection and pipes.
- Utilities like `grep` and `find`.
### Packages (02:18:43)
- Installing, updating, and managing software dependencies and packages.
## Conclusion
The course transforms the command line from a source of intimidation into a powerful tool for productivity and system control.
+111
View File
@@ -0,0 +1,111 @@
---
tags: [neovim, lazyvim, coding-tools, ide, tutorial]
created: 2026-05-29
status: complete
type: lesson
---
# Master Class: In-Depth Guide to LazyVim
LazyVim is not just a configuration; it is a **Neovim setup framework** designed to provide a high-performance, IDE-like experience while remaining modular and easy to customize. It leverages the power of `lazy.nvim` to ensure that your editor starts instantly by loading components only when they are needed.
---
## 1. The Core Philosophy
LazyVim is built on three pillars:
1. **Speed:** Everything is lazy-loaded. If you aren't editing a Python file, the Python LSP doesn't load.
2. **Sane Defaults:** It comes pre-configured with industry-standard settings for UI, indentation, and search.
3. **Modularity:** It separates its core logic from your user configuration, allowing you to update the framework without breaking your personal tweaks.
---
## 2. The Integrated Toolbox
LazyVim integrates several powerful tools into a cohesive workflow:
### A. **lazy.nvim (The Engine)**
The heart of the system. It manages plugin installation, updates, and lazy-loading.
- **Command:** `:Lazy`
- **Capabilities:** Check for updates, profile startup time, and manage plugin states.
### B. **Mason.nvim (The Tool Manager)**
A "package manager" for your external dependencies.
- **Command:** `:Mason`
- **Capabilities:** Easily install and manage LSP servers, DAP (debuggers), linters, and formatters directly from within Neovim.
### C. **nvim-treesitter (The Parser)**
Provides high-performance syntax highlighting and code understanding.
- **Command:** `:TSUpdate`
- **Capabilities:** Better highlighting, indentation, and "incremental selection" (selecting code blocks logically).
### D. **Telescope.nvim / fzf-lua (The Searcher)**
A fuzzy finder that allows you to find anything in your project.
- **Keybinds:** `<leader>ff` (files), `<leader>/` (live grep), `<leader>sk` (keymaps).
---
## 3. Essential Keybindings & Workflow
LazyVim uses the `<Space>` key as the **Leader**. One of its best features is `which-key.nvim`, which displays a popup showing available commands whenever you press your leader key.
### **Navigation**
- `<leader>e`: Toggle **Neo-tree** (File Explorer).
- `H` / `L`: Quickly cycle through open buffers (tabs).
- `<leader>bb`: Switch between open buffers.
- `<leader>fT`: Open a floating terminal.
### **Coding & LSP**
- `K`: Hover documentation (show what a function/variable does).
- `gd`: Go to definition.
- `gr`: Go to references.
- `<leader>ca`: **Code Actions** (Fixes, imports, refactors).
- `<leader>cr`: Rename the symbol under the cursor project-wide.
- `[d` / `]d`: Jump to previous/next error or warning.
### **Git Integration**
- `<leader>gg`: Open **LazyGit** (a full TUI for Git).
- `<leader>gj`: Next Git hunk.
- `<leader>gk`: Previous Git hunk.
---
## 4. Customizing Your Setup
LazyVim's structure is designed to be clean:
- `lua/config/options.lua`: Global Neovim settings (e.g., line numbers, tab widths).
- `lua/config/keymaps.lua`: Your custom keyboard shortcuts.
- `lua/plugins/`: Any `.lua` file created here is automatically loaded as a plugin configuration.
### **Adding a Plugin**
Create `lua/plugins/example.lua`:
```lua
return {
"username/plugin-name",
opts = {
-- plugin configuration goes here
},
}
```
### **Enabling "Extras"**
LazyVim provides "Packs" for specific needs. You can enable them in `lua/config/lazy.lua`:
```lua
require("lazyvim.util").plugin.setup({
spec = {
{ "LazyVim/LazyVim", import = "lazyvim.plugins" },
-- Enable extras like Python, Docker, or Copilot:
{ import = "lazyvim.plugins.extras.lang.python" },
{ import = "lazyvim.plugins.extras.ui.mini-animate" },
{ import = "lazyvim.plugins.extras.coding.copilot" },
{ import = "lua.plugins" },
},
})
```
---
## 5. Why LazyVim?
Unlike building a config from scratch (which can take months to perfect) or using a "thick" distro like LunarVim (which can feel bloated), LazyVim gives you a professional-grade starting point that feels like **your** config. It provides the "glue" that makes LSP, completion, and UI tools work together seamlessly.
---
**Next Steps:**
- Run `:LazyHealth` to check your environment.
- Install `lazygit` on your system to enable the `<leader>gg` shortcut.
- Explore the [Official LazyVim Docs](https://www.lazyvim.org).
+82
View File
@@ -0,0 +1,82 @@
---
tags: [neovim, treesitter, mason, telescope, coding-tools, tutorial]
created: 2026-05-29
status: complete
type: lesson
---
# Deep Dive: The Power Trio (Tree-sitter, Mason, & Telescope)
While LazyVim provides the framework, its "superpowers" come from three specific plugins that redefine how you interact with code. Understanding these tools in-depth will allow you to master your editor.
---
## 1. Tree-sitter: The Syntactic Engine
Traditional editors use "Regex" (Regular Expressions) to highlight code. Regex is just fancy pattern matching. **Tree-sitter** is different: it is a **parser**. It builds a concrete syntax tree (CST) of your source file.
### Why it matters:
- **Semantic Awareness:** It knows the difference between a variable, a function, and a type, even if they have the same name.
- **Incremental Parsing:** It updates the tree instantly as you type, making it incredibly fast.
- **Language-Agnostic:** One engine powers hundreds of languages.
### Key Modules:
- **Highlighting:** Precise, context-aware colors.
- **Incremental Selection:** Logical selection (select word -> select expression -> select function).
- **Indentation:** Uses the tree structure to know exactly where a line should start.
### Pro Commands:
- `:TSInstall <lang>`: Install a specific language parser.
- `:InspectTree`: Opens a side window showing the actual tree structure of your code.
- `:EditQuery`: Advanced tool to create custom highlights or behavior.
---
## 2. Mason.nvim: The Tool Manager
In the past, installing an LSP (Language Server) or a Formatter was a nightmare. You had to use `npm`, `pip`, `go install`, and `cargo` manually. **Mason** abstracts all of this.
### The Architecture:
- **Isolated Binaries:** Mason installs tools in `~/.local/share/nvim/mason/bin/`. It does **not** pollute your system global PATH.
- **The "Bridge":** Mason only *downloads* the tools. Plugins like `mason-lspconfig` are needed to connect those downloads to Neovim's internal LSP client.
### Essential Workflow:
1. Open `:Mason`.
2. Find your tool (e.g., `ruff` for Python, `gopls` for Go).
3. Press `i` to install.
4. Use `ensure_installed` in your config to automate this across machines.
---
## 3. Telescope.nvim: The Central Nervous System
Telescope is a highly extendable fuzzy finder. It is the primary way you find and jump to information.
### The "Picker" Concept:
Everything in Telescope is a **Picker**. A picker consists of three parts:
1. **Finder:** Where the data comes from (files, git, buffers).
2. **Sorter:** How it ranks the results (FZF, native).
3. **Previewer:** Shows you the content before you click.
### Advanced Mappings (Inside Telescope):
- `<C-j>` / `<C-k>`: Move through results.
- `<Tab>`: Select multiple items.
- `<C-q>`: Send all selected items to the **Quickfix List** (powerful for mass refactoring).
- `<C-/>`: Show all keybindings for the current picker.
### Must-Have Extensions:
- **`fzf-native`:** Uses a C-port of FZF for blazing fast sorting.
- **`ui-select`:** Makes *all* Neovim selection menus (like Code Actions) use the Telescope UI.
- **`undo`:** A visual search for your file's history.
---
## How They Work Together
Imagine you are editing a file:
1. **Mason** has installed the `Pyright` LSP server.
2. **Tree-sitter** is providing beautiful syntax highlighting and logical selection.
3. You realize you need to find where a function is used. You trigger **Telescope** (`lsp_references`).
4. Telescope uses the LSP (installed by **Mason**) to find the references and displays them in its UI.
---
**Next Steps:**
- Try `:InspectTree` on a complex file to see how Tree-sitter "sees" your code.
- Run `:Mason` and check if you have any outdated tools.
- Experiment with `live_grep` in Telescope to find text across your entire project.
+20
View File
@@ -0,0 +1,20 @@
---
title: Python Notes Index
created: 2026-05-22
tags:
- python
- index
- moc
category: index
---
# Python Notes Index
All notes under `note/python/`, grouped by subfolder.
```dataview
LIST rows.file.link
FROM "note"
WHERE file.name != "index"
GROUP BY file.folder
SORT file.folder ASC
```
@@ -0,0 +1,176 @@
# SQLAlchemy 2.0 Tutorial (Recipe Scraper Edition)
This guide covers SQLAlchemy 2.0, the industry-standard SQL toolkit and Object-Relational Mapper (ORM) for Python. We'll use the **Recipe Web Scraper** project models as our primary examples.
---
## 1. What is SQLAlchemy?
SQLAlchemy has two main components:
1. **Core**: A SQL abstraction layer (SQL Expression Language, Schema definitions, Engine).
2. **ORM**: A layer on top of Core that maps Python classes to database tables.
In this project, we primarily use the **ORM** to treat recipes and ingredients as Python objects.
---
## 2. Defining Models (The Modern Way)
SQLAlchemy 2.0 introduced a type-hint-centric way to define models using `Mapped` and `mapped_column`.
### The Base Class
All models inherit from a common `Base` class created from `DeclarativeBase`.
```python
from sqlalchemy.orm import DeclarativeBase
class Base(DeclarativeBase):
pass
```
### Example: The Recipe Model
```python
from sqlalchemy import String, Integer, DateTime
from sqlalchemy.orm import Mapped, mapped_column, relationship
from datetime import datetime
class Recipe(Base):
__tablename__ = "recipes" # Name of the table in the DB
# Primary Key
id: Mapped[int] = mapped_column(primary_key=True)
# Simple Columns (SQLAlchemy infers types from Mapped[T])
url: Mapped[str] = mapped_column(String, unique=True, index=True)
title: Mapped[str]
total_time: Mapped[int | None] # Optional column (nullable=True)
# Column with a default value
scraped_at: Mapped[datetime] = mapped_column(DateTime, default=datetime.utcnow)
# Relationships (Defined in section 4)
ingredients: Mapped[list["Ingredient"]] = relationship(back_populates="recipe")
```
---
## 3. Engine and Session
### The Engine
The **Engine** is the starting point for any SQLAlchemy application. It manages a pool of connections to the database.
```python
from sqlalchemy import create_engine
# SQLite: The '///' means relative path to the current directory
engine = create_engine("sqlite:///recipes.db", echo=True)
# echo=True logs all SQL commands to the terminal (great for debugging)
```
### Creating Tables
You can tell SQLAlchemy to create all tables defined in your models:
```python
Base.metadata.create_all(engine)
```
### The Session
The **Session** handles the conversation with the database. Use `sessionmaker` to create a factory for sessions.
```python
from sqlalchemy.orm import sessionmaker
SessionLocal = sessionmaker(bind=engine)
# Use as a context manager to ensure the connection is closed
with SessionLocal() as session:
# do work here
pass
```
---
## 4. Relationships (1-to-Many)
In our project, one `Recipe` has many `Ingredients`.
### Foreign Key
The "child" table (`Ingredient`) must have a column pointing to the "parent" table (`Recipe`).
```python
from sqlalchemy import ForeignKey
class Ingredient(Base):
__tablename__ = "ingredients"
id: Mapped[int] = mapped_column(primary_key=True)
# Links to 'recipes.id'
recipe_id: Mapped[int] = mapped_column(ForeignKey("recipes.id"))
text: Mapped[str]
# Back-reference to the parent Recipe object
recipe: Mapped["Recipe"] = relationship(back_populates="ingredients")
```
### Cascades
`cascade="all, delete-orphan"` ensures that if you delete a Recipe, all its Ingredients are also deleted automatically.
---
## 5. CRUD Operations
### Create (Insert)
```python
with SessionLocal() as session:
new_recipe = Recipe(title="Pasta Carbonara", url="https://example.com/pasta")
session.add(new_recipe)
session.commit() # Save to DB
```
### Read (Select)
```python
from sqlalchemy import select
with SessionLocal() as session:
# 1. Get by ID
recipe = session.get(Recipe, 1)
# 2. Filter by column
stmt = select(Recipe).where(Recipe.title == "Pasta Carbonara")
result = session.execute(stmt).scalars().first()
# 3. Get all
all_recipes = session.query(Recipe).all() # Older syntax, still common
```
### Update
```python
with SessionLocal() as session:
recipe = session.get(Recipe, 1)
recipe.title = "Authentic Pasta Carbonara"
session.commit()
```
### Delete
```python
with SessionLocal() as session:
recipe = session.get(Recipe, 1)
session.delete(recipe)
session.commit()
```
---
## 6. Common Pitfalls
1. **Lazy Loading**: By default, SQLAlchemy doesn't load relationships until you access them. This can cause "N+1" performance issues. Use `joinedload` to fetch everything in one query.
2. **Session Lifecycle**: Always use a context manager (`with session:`) or close your sessions manually.
3. **Commit vs Flush**: `session.flush()` sends changes to the DB but doesn't permanentize them. `session.commit()` makes them permanent.
---
## 7. Next Steps: Migrations with Alembic
As your models change (e.g., you add a `rating` column), you shouldn't just delete the DB and start over. **Alembic** is the tool used to handle database migrations.
Install it with: `pip install alembic`
+93
View File
@@ -0,0 +1,93 @@
---
title: Poetry Setup and Project Initialization Guide
created: 2026-05-22
tags:
- python
- python_tool
- poetry
- guide
category: python_tool
status: reference
up: "[[project/index]]"
related:
- "[[poetry_project_ideas]]"
- "[[__init__.py explained]]"
source:
author:
published:
---
# Poetry Setup and Project Initialization Guide
This guide explains how to install Poetry and start a new Python project, based on the concepts from "Introduction to Poetry - Python Dependency Management".
## 1. How to Install Poetry
While the introductory notes focus on usage, the standard way to install Poetry is via the official installer script.
### macOS / Linux / WSL
Open your terminal and run:
```bash
curl -sSL https://install.python-poetry.org | python3 -
```
### Windows (PowerShell)
Open PowerShell and run:
```powershell
(Invoke-WebRequest -Uri https://install.python-poetry.org -UseBasicParsing).Content | py -
```
### Verification
After installation, restart your terminal and verify by running:
```bash
poetry --version
```
---
## 2. How to Start a Project
There are two main ways to start a project with Poetry:
### Method A: Creating a New Project (Recommended for new folders)
To create a new project with a predefined folder structure:
```bash
poetry new my-project
```
This creates a directory named `my-project` with the following structure:
```text
my-project/
├── pyproject.toml
├── README.md
├── my_project/
│ └── __init__.py
└── tests/
└── __init__.py
```
### Method B: Initializing an Existing Project
If you already have a project folder and want to add Poetry to it:
1. Navigate to your project directory:
```bash
cd my-existing-project
```
2. Run the interactive initialization command mentioned in the introduction:
```bash
poetry init
```
This will walk you through creating your `pyproject.toml` file interactively.
---
## 3. Key Concepts from the Introduction
- **`pyproject.toml`**: The single source of truth for your project configuration (replaces `requirements.txt`, `setup.py`, etc.).
- **Deterministic Resolution**: Poetry ensures your dependencies are resolved correctly using a lockfile (`poetry.lock`).
- **Isolation**: Poetry automatically manages virtual environments for you, ensuring your global Python installation stays clean.
## 4. Basic Workflow Commands
Once your project is started, use these commands to manage it:
- `poetry add <package>`: Add and install a new dependency.
- `poetry install`: Install all dependencies defined in `pyproject.toml`.
- `poetry shell`: Activate the project's virtual environment.
- `poetry run <command>`: Run a command inside the virtual environment without activating it.
@@ -0,0 +1,152 @@
---
title: What is __init__.py?
created: 2026-05-22
tags:
- python
- python_note
- packaging
- reference
category: python_note
status: reference
up: "[[project/index]]"
related:
- "[[poetry_guide]]"
- "[[poetry_project_ideas]]"
source:
author:
published:
---
# What is `__init__.py`?
`__init__.py` is the file that tells Python "this folder is a package". When you put it inside a directory, that directory becomes importable like a module.
## The Core Purpose
Without `__init__.py` (in older Python, ≤3.2), a folder was just a folder — Python could not `import` from it. With it, the folder becomes a **package** you can do this with:
```python
from recipe_scraper.scraper import fetch_recipe
```
Here, `recipe_scraper/` is a package because it contains `__init__.py`.
> Note: Since Python 3.3, "namespace packages" allow imports without `__init__.py`, but **regular packages still use it** because it gives you more control (init code, explicit exports, IDE/tooling support).
---
## What It Does in Practice
### 1. Marks the folder as a package
Even an **empty** `__init__.py` is meaningful. It signals to Python:
> "Treat this directory as something you can import from."
### 2. Runs initialization code
Anything inside `__init__.py` runs **the first time** the package is imported. This is useful for:
- Setting up logging
- Loading config
- Registering plugins
```python
# recipe_scraper/__init__.py
import logging
logging.getLogger(__name__).addHandler(logging.NullHandler())
```
### 3. Controls the package's public API
You can re-export things so users don't need to know the internal file layout:
```python
# recipe_scraper/__init__.py
from .scraper import fetch_recipe
from .parser import parse_recipe
from .db import save_recipe
__all__ = ["fetch_recipe", "parse_recipe", "save_recipe"]
```
Now consumers can write the short form:
```python
from recipe_scraper import fetch_recipe # clean
# instead of:
from recipe_scraper.scraper import fetch_recipe # verbose
```
### 4. Defines package metadata
A common pattern is exposing a version string:
```python
# recipe_scraper/__init__.py
__version__ = "0.1.0"
```
---
## How It Fits the Recipe Scraper Project
Given the folder structure from [[Recipe Web Scraper - Getting Started]]:
```
recipe-scraper/
├── pyproject.toml
└── recipe_scraper/
├── __init__.py ← makes this a package
├── scraper.py
├── parser.py
├── db.py
└── cli.py
```
The `__init__.py` here lets you:
1. Run `poetry run python -m recipe_scraper.cli` — only works because `recipe_scraper` is a package.
2. Import cleanly from anywhere in the project:
```python
from recipe_scraper.db import Recipe
```
3. Optionally expose a tidy top-level API:
```python
# recipe_scraper/__init__.py
from .scraper import scrape_url
from .db import Recipe, init_db
__version__ = "0.1.0"
__all__ = ["scrape_url", "Recipe", "init_db"]
```
So users of your library can just do:
```python
import recipe_scraper
recipe_scraper.scrape_url("https://...")
```
---
## When to Leave It Empty vs. Fill It
| Situation | What to put in `__init__.py` |
| :--- | :--- |
| Internal-only package, no public API | Empty file |
| You want a clean import surface | Re-exports + `__all__` |
| Library shipped to PyPI | `__version__`, re-exports, maybe logging setup |
| One-time setup needed (config, env) | Initialization code at the top |
**Rule of thumb:** start with an empty `__init__.py`. Only add code when you have a concrete reason — re-exports, version, or setup. Don't put heavy logic in `__init__.py`; it runs on every import.
---
## Common Gotchas
- **Heavy imports slow everything down.** If `__init__.py` imports a big library (e.g., `pandas`), every `import recipe_scraper.anything` pays that cost. Keep it light.
- **Circular imports** often start in `__init__.py`. If `__init__.py` imports from `scraper.py`, and `scraper.py` imports from the package root, you get a cycle. Use lazy imports or restructure.
- **Tests need it too.** A `tests/` folder usually has an empty `__init__.py` so pytest can discover test modules consistently (though pytest's `rootdir` config can avoid this).
---
## TL;DR
- `__init__.py` = "this folder is a Python package".
- Can be empty — just its presence matters.
- Use it to **re-export** the public API, set `__version__`, or run small setup.
- Keep it light: every import of the package runs it.
## Related
- [[Recipe Web Scraper - Getting Started]]
- [[poetry_guide]] — Poetry's `poetry new` creates this file automatically
@@ -0,0 +1,177 @@
---
title: __init__.py vs main.py
created: 2026-05-22
tags:
- python
- python_note
- packaging
- reference
category: python_note
status: reference
up: "[[project/index]]"
related:
- "[[__init__.py explained]]"
- "[[Recipe Web Scraper - Getting Started]]"
source:
author:
published:
---
# `__init__.py` vs `main.py`
Both files live inside a Python package, but they have very different jobs. This note uses the [[Recipe Web Scraper - Getting Started|recipe scraper]] project as the running example.
## The Recipe Scraper Layout
```
recipe-scraper/
└── src/
└── recipe_scraper/
├── __init__.py ← marks this as a package
├── main.py ← CLI entry point (you run this)
├── scraper.py
├── models.py
├── db.py
└── exporters.py
```
Two files, two completely different purposes.
---
## At a Glance
| Aspect | `__init__.py` | `main.py` |
| :--- | :--- | :--- |
| **Purpose** | Marks the folder as a package; sets up the package | Holds the **program's entry point** (the code you run) |
| **Required?** | Yes (for regular packages) | No — it's a convention, name can be anything |
| **When it runs** | Automatically, on **any** import of the package | Only when you **explicitly run it** |
| **Typical contents** | Re-exports, `__version__`, light setup | `main()` function, CLI parsing, `if __name__ == "__main__"` |
| **Who calls it** | Python's import system | The user (via `python -m`, `poetry run`, etc.) |
| **Should be lightweight?** | **Yes** — runs every import | Can be heavy — runs once when invoked |
---
## `__init__.py` — The Package Setup File
Its job is to tell Python "this folder is a package" and optionally configure what the package looks like from the outside.
```python
# src/recipe_scraper/__init__.py
from .scraper import scrape_url
from .db import Recipe, init_db
__version__ = "0.1.0"
__all__ = ["scrape_url", "Recipe", "init_db"]
```
After this, anyone using the library can write:
```python
from recipe_scraper import scrape_url
scrape_url("https://example.com/recipe")
```
instead of digging into `recipe_scraper.scraper`. **It never gets "run" directly — it runs implicitly whenever the package is imported.**
See [[__init__.py explained]] for the full picture.
---
## `main.py` — The Program Entry Point
Its job is to be the **code that actually executes when you launch the program**. For the recipe scraper, it's where the CLI lives:
```python
# src/recipe_scraper/main.py
import typer
from .scraper import scrape_url
from .db import init_db
app = typer.Typer()
@app.command()
def scrape(url: str):
"""Scrape a recipe URL and save it to the database."""
init_db()
recipe = scrape_url(url)
print(f"Saved: {recipe.title}")
@app.command()
def list_recipes():
"""List all saved recipes."""
...
def main():
app()
if __name__ == "__main__":
main()
```
You run it like:
```bash
poetry run python -m recipe_scraper.main scrape https://example.com/recipe
```
Or, if `pyproject.toml` defines an entry point pointing at `main:main`, just:
```bash
poetry run recipe-scraper scrape https://example.com/recipe
```
---
## The Key Mental Model
> `__init__.py` answers: **"What is this package?"**
> `main.py` answers: **"What happens when you run this program?"**
- Importing the package → `__init__.py` runs.
- Running the program → `main.py` runs (and it imports things, which triggers `__init__.py`).
So in a typical execution:
```
$ poetry run python -m recipe_scraper.main scrape ...
│
├── Python loads `recipe_scraper` package
│ └── runs __init__.py (sets up exports, version, etc.)
│
└── Python runs main.py as the module
└── parses CLI args, calls scrape_url(), saves to DB
```
`__init__.py` runs **first** (because `main.py` is inside the package), but it does almost nothing visible. `main.py` is where the actual work happens.
---
## Common Pitfalls
### 1. Putting CLI code in `__init__.py`
**Don't.** It will run every time *anything* imports your package, including your tests. Keep entry-point code in `main.py` (or `cli.py`).
### 2. Heavy imports in `__init__.py`
If `__init__.py` imports `pandas`, `torch`, or anything slow, every import of your package pays that cost. For the recipe scraper, avoid importing `recipe-scrapers` or SQLAlchemy at the top of `__init__.py` — import them inside the modules that need them.
### 3. Forgetting `if __name__ == "__main__":` in `main.py`
Without that guard, `main.py`'s code runs whenever the module is **imported**, not just when it's **executed**. Always wrap the entry-point call:
```python
if __name__ == "__main__":
main()
```
### 4. Confusing `main.py` with `__main__.py`
- `main.py` — just a convention; you have to point at it explicitly (`python -m recipe_scraper.main`).
- `__main__.py` — a **special name**: lets you run `python -m recipe_scraper` (no `.main` needed).
If you want `python -m recipe_scraper` to work, rename `main.py` to `__main__.py`. Many projects keep both: `__main__.py` is a one-liner that calls into `main.py`.
---
## TL;DR
- **`__init__.py`** = package marker + public-API shaping. Runs on every import. Keep it light.
- **`main.py`** = program entry point. Runs when you invoke the CLI. This is where the action is.
- They're complementary, not alternatives. In the recipe scraper, `__init__.py` exposes the library; `main.py` is the CLI you actually run.
## Related
- [[__init__.py explained]]
- [[Recipe Web Scraper - Getting Started]]
- [[poetry_guide]]