Document Pipeline is a Phoenix LiveView application for managing construction
loan projects and the documents (invoices, budgets, change orders, pay
applications) submitted against them. Uploaded documents are classified and
have structured fields/line items extracted from them via a pluggable AI
pipeline (Claude by default), processed asynchronously with Oban, and
reviewed/corrected live in the UI over Phoenix.PubSub — no separate API
layer, everything is server-rendered LiveView.
- Elixir
~> 1.15and a compatible Erlang/OTP - PostgreSQL running locally (defaults to
localhost, seeconfig/dev.exs/config/test.exs) - Node is not required — assets are built with
esbuild/tailwind, whichmix setupinstalls automatically
# Install dependencies, create/migrate the dev database, and build assets
mix setupDev and test both default to the mock classifier/extractor adapters
(config/dev.exs, config/test.exs), so the full upload → classify →
extract → correct loop works locally with no API key required — that's
the intended way to run and demo this app day to day.
To exercise the real Anthropic-backed adapters instead:
# config/dev.exs
config :document_pipeline,
classifier_adapter: DocumentPipeline.AI.Classifier.DefaultAdapter,
extractor_adapter: DocumentPipeline.AI.Extractor.DefaultAdapterexport ANTHROPIC_API_KEY=sk-ant-...config/runtime.exs reads this into :document_pipeline, :anthropic_api_key
at boot (dev and prod alike), which is what the default adapters call
Application.get_env(:document_pipeline, :anthropic_api_key) to fetch.
Production additionally requires DATABASE_URL and SECRET_KEY_BASE — see
config/runtime.exs for the full list and defaults.
iex -S mix phx.serverThen visit:
localhost:4000/workspace— the main screen: upload a document, watch it classify/extract live, correct its type, hand-edit fieldslocalhost:4000/projects,/documents,/document_fields,/document_line_items— plain CRUD screens (mix phx.gen.livescaffolding) for poking at the underlying data directly; not linked from the main nav
mix testDocumentPipeline.Projects/DocumentPipeline.Documents— the core contexts. Projects own documents; documents own extracted fields and line items.DocumentPipeline.AI.Classifier/DocumentPipeline.AI.Extractor— behaviours with pluggable adapters (a real Claude-backed adapter and a mock adapter with deterministic fixture data), selected via:document_pipeline, :classifier_adapter/:extractor_adapter.DocumentPipeline.Workers.ProcessDocumentWorker— an Oban worker that runs an uploaded document through text extraction, classification, and field/line-item extraction, updating itsstatusthroughout and broadcasting completion overPhoenix.PubSubon the"document#{id}"topic.DocumentPipelineWeb.WorkspaceLive— the primary UI. Subscribes to each visible document's PubSub topic so extraction results and status changes appear without a page refresh; also drives type correction and inline field editing.
erDiagram
PROJECT ||--o{ DOCUMENT : "has_many :documents"
DOCUMENT ||--o{ DOCUMENT_FIELD : "has_many :document_fields"
DOCUMENT ||--o{ DOCUMENT_LINE_ITEM : "has_many :document_line_items"
PROJECT {
bigint id PK
string name
string address
decimal total_budget
string status "free text, not enum-constrained"
}
DOCUMENT {
bigint id PK
bigint project_id FK
string filename
string content_type
string storage_path
integer file_size
string domain_type "invoice | budget | change_order | pay_application"
string domain_type_source "system | user"
string status "uploaded | processing | processed | failed"
text raw_text
string error_message
}
DOCUMENT_FIELD {
bigint id PK
bigint document_id FK
string field_name
string field_value
float confidence
string source "system | user"
}
DOCUMENT_LINE_ITEM {
bigint id PK
bigint document_id FK
string description
decimal amount
string category
integer line_number
string source "system | user"
}
Walkthrough:
- A Project is created (e.g.
"Montana - Phase 3") and owns any number of Documents uploaded against it (Project.has_many :documents/Document.belongs_to :project, FKdocument.project_id). - Uploading a document (
Documents.upload_document/4) writes the file topriv/uploads/<uuid>_<filename>, inserts aDocumentrow withstatus: "uploaded", and enqueuesProcessDocumentWorkervia Oban. - The worker walks the document through
statusvalues in order —uploaded → processing → processedon success, orfailed(witherror_messageset to the failure reason) if any step errors. - Once classified, the extractor produces the document's DocumentFields
(key/value pairs like
vendor_name,invoice_number) and DocumentLineItems (cost-breakdown rows) using a schema specific to the domain type —change_ordernever has line items, the other three types do. - Both
DocumentFieldandDocumentLineItemcarry asourcecolumn ("system"vs"user"): AI-extracted fields start as"system"with aconfidencescore; hand-correcting one viaupdate_field/2flips it to"user"withconfidence: 1.0. Correcting a document's type viacorrect_document_type/2deletes all of its fields/line items and re-triggers extraction against the corrected type, stampingdomain_type_source: "user".