Imp Brings DSPy's Declarative AI Programming Model to Elixir on the BEAM

Imp Brings DSPy's Declarative AI Programming Model to Elixir on the BEAM

Stackademic

Imp ports DSPy's declarative AI programming to Elixir/BEAM with signatures, optimizers, RAG, and supervision-tree deployment.

A New Language for AI Systems Lands on the BEAM

On October 4, 2026, the Elixir community gained a library that could reshape how developers build production AI applications on the BEAM virtual machine. Imp, a full port of Stanford's influential DSPy framework, brings declarative AI programming to Elixir with native support for supervision trees, concurrency primitives, and the fault-tolerance philosophy that defines OTP applications.

Where most teams still assemble LLM features from ad hoc prompt strings, retry logic scattered across controllers, and evaluation notebooks that never quite sync with production, Imp offers a structured programming model: define what a module should accomplish, compose modules into pipelines and agents, optimize prompts and weights systematically, and run the result as first-class Elixir processes.

For teams already running real-time systems, API gateways, and event-driven architectures on Elixir, the release answers a practical question: does declarative AI belong in the same language as your orchestration layer? Imp's authors argue yes — and early adopters are testing that claim in RAG services, agent workflows, and multi-step reasoning pipelines.

Why DSPy Mattered — and Why Elixir Needs Its Own Port

DSPy emerged from research demonstrating that treating LLM applications as programs rather than prompts improves reliability, maintainability, and optimization outcomes. Instead of manually tweaking wording, developers define signatures — typed input-output contracts — and let frameworks handle prompt synthesis, demonstration selection, and metric-driven improvement.

Python's DSPy ecosystem grew quickly because it aligned with notebook-centric experimentation and PyTorch-adjacent ML culture. But many production systems — especially those handling websockets, distributed cron, stream processing, and telephony-scale concurrency — already live on Elixir and Erlang. Those teams faced a fork: bridge to Python microservices with operational overhead, or reimplement DSPy concepts by hand with inconsistent patterns.

Imp eliminates that false choice by porting the conceptual core: signatures, modules, teleprompters (optimizers), retrieval integration, and composable forward passes — implemented idiomatically for Elixir.

Declarative Signatures: Contracts Instead of Prompt Fragments

At the center of Imp is the signature, a declarative specification of a transformation. A signature names fields, types, and instructional intent without embedding final prompt prose in application code. The framework materializes prompts from signatures at compile or runtime depending on configuration, enabling centralized updates when models or policies change.

Consider a support ticket classifier. Imperative style hardcodes:

"Classify the following ticket into billing, technical, or sales. Ticket: ..."

Imp style defines a signature module:

defmodule Support.TicketClassifier do
  use Imp.Signature

  signature do
    input :ticket_text, :string
    output :category, :enum, values: [:billing, :technical, :sales]
    instruction "Classify customer support tickets into the correct department."
  end
end

The application calls the module; Imp handles prompt assembly, parser validation, and structured output coercion. When the business adds a fraud category, you update the signature — not forty string templates across the codebase.

This is more than ergonomics. Signatures become unit boundaries for testing, telemetry, and optimization.

Modules, Composition, and the Forward Pass Mental Model

Imp modules encapsulate logic that may invoke models, tools, or other modules. DSPy's forward abstraction — a predictable execution path with explicit inputs and outputs — maps cleanly onto Elixir functions and with pipelines.

Developers compose modules into graphs representing multi-step reasoning: retrieve context, draft an answer, verify citations, refine tone. Because composition is code, teams apply familiar refactors — extract module, inject dependency, mock downstream model — instead of copying prompt blocks.

Imp also supports conditional branches and lightweight agent loops where a module decides whether to call tools, request more retrieval, or terminate. The agent pattern is notoriously easy to prototype in notebooks and notoriously hard to operate in production. Imp's bet is that BEAM process isolation plus explicit module boundaries tame runaway loops better than unstructured scripts.

DSPy's breakthrough for many teams was teleprompting — optimization procedures that search over instructions, few-shot examples, and module parameters using a developer-provided metric. Imp ports this optimizer stack, enabling workflows like:

  • Bootstrap few-shot: automatically select demonstrations that improve validation accuracy.
  • Instruction search: explore paraphrased instructions ranked by metric scores.
  • Multi-module coordination: optimize pipelines where downstream modules consume upstream outputs.

In Elixir, optimizers run as supervised tasks. You can schedule re-optimization when evaluation sets grow, when models version, or when product metrics drift — without taking production nodes offline.

Example metric functions look like ordinary Elixir:

def metric(prediction, expected) do
  if prediction.category == expected.category, do: 1.0, else: 0.0
end

Imp feeds predictions and gold labels through the metric during optimizer runs, accumulating scores that guide search. This closes the loop between offline evaluation and deployed behavior — a loop that often breaks when prompts live only in Slack threads.

RAG as a First-Class Pipeline Concern

Retrieval-augmented generation in Imp is not a bolt-on tutorial. Vector stores, chunkers, rerankers, and context assemblers integrate as modules in the same composition graph as generators. Signatures express what retrieved context must supply — citations, date bounds, product SKUs — and optimizers can tune retrieval queries and formatting instructions jointly.

For BEAM deployments, retrieval workloads benefit from concurrent fan-out: query multiple indices, merge results, deduplicate, and pass ranked context to generation processes under timeout budgets. Elixir's strength in coordinating IO-bound work maps naturally onto RAG hot paths that stall Python services when implemented naively.

Model-Agnostic Design in a Multi-Provider World

Imp does not lock teams to a single vendor API. The library ships with adapters for major hosted models and local inference servers, using a common interface for chat completions, embeddings, and tool calls. Configuration swaps models per module — useful when cheap models handle classification and frontier models handle synthesis.

Model-agnosticism is operational, not ideological. Production systems routinely route traffic based on cost, latency, and jurisdiction. Imp modules carry model preferences in metadata, enabling dynamic routing policies enforced at runtime by orchestrator processes.

Running in Supervision Trees: OTP as an AI Reliability Layer

The distinguishing architectural claim of Imp is not merely DSPy on Elixir syntax — it is DSPy inside OTP. LLM calls run as supervised processes with restart strategies, circuit breakers, and backpressure compatible with existing Phoenix and Broadway applications.

When a model endpoint spikes latency, Imp supervision can shed load, fail over to degraded modules, or return structured errors to callers — behaviors that are tedious to recreate in one-off Python scripts but routine in OTP systems. Telemetry hooks emit events compatible with OpenTelemetry stacks, so AI modules appear in the same dashboards as database and HTTP spans.

For SRE teams allergic to "another AI microservice," Imp offers a consolidation story: the agents live where the rest of the app lives, governed by the same deployment, secret management, and on-call playbooks.

Imp versus Imperative Prompt Engineering

Imperative prompt engineering optimizes for demo speed. It excels in hackathons and proof-of-concepts. It scales poorly when:

  • Multiple engineers mutate the same prompts without version discipline.
  • Evaluation metrics exist but do not propagate back to prompt changes.
  • Retrieval, tools, and generation evolve on different cadences.
  • Incidents require knowing exactly which prompt version served a bad output.

Imp trades some initial learning curve for maintainability. Signatures are reviewable in pull requests. Optimizer runs are reproducible artifacts. Module graphs document system intent more clearly than a prompts folder.

The comparison is not absolute. Imp does not remove the need for domain judgment or careful dataset curation. It channels that judgment into structures optimizers and tests can act upon.

Getting Started: A Minimal Agent Loop

A typical first project implements a research assistant:

  1. Retriever module fetches documents from an internal index.
  2. Synthesizer module signature requires cited bullet summaries.
  3. Verifier module checks that each bullet references retrieved IDs.
  4. Supervisor manages retries if verification fails.

Each step is a module with explicit inputs and outputs. The agent loop runs until verification passes or attempts exhaust. Tests stub modules with deterministic outputs — no network calls in CI.

Imp's documentation emphasizes incremental adoption: wrap one imperative prompt as a signature, measure, optimize, then expand composition.

Ecosystem Fit: Phoenix, LiveView, and Broadway

Early integrators report pairing Imp with:

  • Phoenix for HTTP and channel interfaces to agent workflows.
  • LiveView for human-in-the-loop review of module outputs before commit actions.
  • Broadway for pipeline ingestion where each message triggers a supervised Imp graph.

These pairings highlight a design goal: AI features should feel like Elixir features, not foreign attachments.

Limitations and Honest Tradeoffs

Imp is new. Ports of research frameworks rarely achieve one-to-one feature parity on day one. Teams should expect:

  • Optimizer runtimes that are compute-intensive and should run off the request path.
  • Learning curve for developers unfamiliar with DSPy concepts.
  • Ecosystem gaps compared to Python's ML package breadth for exotic model types.

Imp is also not a replacement for data engineering. Garbage retrieval still produces garbage answers, no matter how elegant the signatures.

Who Should Evaluate Imp Now?

Strong fit:

  • Elixir shops adding LLM features without spinning up Python ops silos.
  • Teams frustrated by prompt sprawl and seeking metric-driven optimization.
  • Systems requiring concurrent RAG and strict supervision policies.

Weaker fit:

  • Teams heavily invested in PyTorch fine-tuning pipelines with no BEAM footprint.
  • One-off demos where a single prompt in a script suffices.

The Bigger Picture for Declarative AI

Imp's release on October 4, 2026, arrives as the industry acknowledges prompt engineering alone does not scale to dependable software. Frameworks like DSPy popularized programs over strings. Imp tests whether that philosophy resonates in a concurrency-first, fault-tolerant runtime — not only in research Python.

If declarative AI programming becomes the default abstraction layer above raw model APIs, ports like Imp matter because production AI will not live in a single language. It will live where companies already run reliable systems. For a growing cohort of teams, that place is the BEAM.

The library is available now for experimentation. The meaningful benchmark is not whether Imp wins GitHub stars on launch day, but whether Elixir teams ship more testable, optimizable, operable AI modules six months from now — with fewer emergency prompt rollbacks and fewer mystery behaviors in production logs.

That is the promise of declarative AI on the BEAM. Imp is the first full-throated attempt to deliver it.