Files
framework_note/note/python/python_note/__init__.py vs main.py.md
T
2026-05-30 16:02:33 -04:00

5.5 KiB

title, created, tags, category, status, up, related, source, author, published
title created tags category status up related source author published
__init__.py vs main.py 2026-05-22
python
python_note
packaging
reference
python_note reference project/index
__init__.py explained
Recipe Web Scraper - Getting Started

__init__.py vs main.py

Both files live inside a Python package, but they have very different jobs. This note uses the Recipe Web Scraper - Getting Started project as the running example.

The Recipe Scraper Layout

recipe-scraper/
└── src/
    └── recipe_scraper/
        ├── __init__.py     ← marks this as a package
        ├── main.py         ← CLI entry point (you run this)
        ├── scraper.py
        ├── models.py
        ├── db.py
        └── exporters.py

Two files, two completely different purposes.


At a Glance

Aspect __init__.py main.py
Purpose Marks the folder as a package; sets up the package Holds the program's entry point (the code you run)
Required? Yes (for regular packages) No — it's a convention, name can be anything
When it runs Automatically, on any import of the package Only when you explicitly run it
Typical contents Re-exports, __version__, light setup main() function, CLI parsing, if __name__ == "__main__"
Who calls it Python's import system The user (via python -m, poetry run, etc.)
Should be lightweight? Yes — runs every import Can be heavy — runs once when invoked

__init__.py — The Package Setup File

Its job is to tell Python "this folder is a package" and optionally configure what the package looks like from the outside.

# src/recipe_scraper/__init__.py
from .scraper import scrape_url
from .db import Recipe, init_db

__version__ = "0.1.0"
__all__ = ["scrape_url", "Recipe", "init_db"]

After this, anyone using the library can write:

from recipe_scraper import scrape_url
scrape_url("https://example.com/recipe")

instead of digging into recipe_scraper.scraper. It never gets "run" directly — it runs implicitly whenever the package is imported.

See __init__.py explained for the full picture.


main.py — The Program Entry Point

Its job is to be the code that actually executes when you launch the program. For the recipe scraper, it's where the CLI lives:

# src/recipe_scraper/main.py
import typer
from .scraper import scrape_url
from .db import init_db

app = typer.Typer()

@app.command()
def scrape(url: str):
    """Scrape a recipe URL and save it to the database."""
    init_db()
    recipe = scrape_url(url)
    print(f"Saved: {recipe.title}")

@app.command()
def list_recipes():
    """List all saved recipes."""
    ...

def main():
    app()

if __name__ == "__main__":
    main()

You run it like:

poetry run python -m recipe_scraper.main scrape https://example.com/recipe

Or, if pyproject.toml defines an entry point pointing at main:main, just:

poetry run recipe-scraper scrape https://example.com/recipe

The Key Mental Model

__init__.py answers: "What is this package?" main.py answers: "What happens when you run this program?"

  • Importing the package → __init__.py runs.
  • Running the program → main.py runs (and it imports things, which triggers __init__.py).

So in a typical execution:

$ poetry run python -m recipe_scraper.main scrape ...
        │
        ├── Python loads `recipe_scraper` package
        │     └── runs __init__.py  (sets up exports, version, etc.)
        │
        └── Python runs main.py as the module
              └── parses CLI args, calls scrape_url(), saves to DB

__init__.py runs first (because main.py is inside the package), but it does almost nothing visible. main.py is where the actual work happens.


Common Pitfalls

1. Putting CLI code in __init__.py

Don't. It will run every time anything imports your package, including your tests. Keep entry-point code in main.py (or cli.py).

2. Heavy imports in __init__.py

If __init__.py imports pandas, torch, or anything slow, every import of your package pays that cost. For the recipe scraper, avoid importing recipe-scrapers or SQLAlchemy at the top of __init__.py — import them inside the modules that need them.

3. Forgetting if __name__ == "__main__": in main.py

Without that guard, main.py's code runs whenever the module is imported, not just when it's executed. Always wrap the entry-point call:

if __name__ == "__main__":
    main()

4. Confusing main.py with __main__.py

  • main.py — just a convention; you have to point at it explicitly (python -m recipe_scraper.main).
  • __main__.py — a special name: lets you run python -m recipe_scraper (no .main needed).

If you want python -m recipe_scraper to work, rename main.py to __main__.py. Many projects keep both: __main__.py is a one-liner that calls into main.py.


TL;DR

  • __init__.py = package marker + public-API shaping. Runs on every import. Keep it light.
  • main.py = program entry point. Runs when you invoke the CLI. This is where the action is.
  • They're complementary, not alternatives. In the recipe scraper, __init__.py exposes the library; main.py is the CLI you actually run.