etl4s

Your First Pipeline

In etl4s, everything is either:

Nodes are simple lambdas from In => Out

import etl4s._

val double = Node[Int, Int](_ * 2)

double(5) // 10

You can run them like functions or be more deliberate with .unsafeRun(In) ... just a matter of taste.

Create new nodes by chaining existing ones together:

val pipeline = 
     double ~> double

pipeline(5) // 20

Here is a more substantive example:

val extract5    = Node(5)
val timesTwo    = Node[Int, Int](_ * 2)
val consoleLoad = Node[Int, Unit](x => println(s"Result: $x"))

val p =
     extract5 ~> timesTwo ~> consoleLoad

p.unsafeRun()

This will give:

Result: 10

To improve readability and express intent, etl4s defines three aliases: Extract, Transform and Load. All behave the same under the hood.

type Extract[In, Out]   = Node[In, Out]
type Transform[In, Out] = Node[In, Out]
type Load[In, Out]      = Node[In, Out]

You can use other operators like & to fan out and stitch your graphs

val double  = Node[Int, Int](_ * 2)
val triple  = Node[Int, Int](_ * 3)
val combine = Node[(Int, Int), Int] { case (a, b) => a + b }

val p =
     extract5 ~> (double & triple) ~> combine

p.unsafeRun() // 25

One of the key benefits of etl4s is that you can separate configuration (the "knobs to turn") from your actual flow of data.

val YEAR = 2025

val loadData = Node[Any, String] { _ =>
  println(s"Loading $YEAR data")
  "TEST DATA"
}

loadData.unsafeRun()

Prints Loading 2025 data and returns "TEST DATA".

That works for one value, but it doesn't compose: every node that reaches for year is an invisible, untyped dependency.

etl4s makes the knob an explicit input instead, declare it with .requires, then .provide it once at the edge:

case class Config(year: Int)

val loadData = Node[Unit, String].requires[Config] { 
    config => _ => s"Loading ${config.year} data"
}

loadData.provide(Config(2025)).unsafeRun(()) // "Loading 2025 data"

.requires turns the node into a Reader[Config, Node[...]], and config-aware and plain nodes compose together with the same ~>. See Configuration for details.

Pipelines are values

Building a pipeline runs nothing. Every combinator (~>, &, >>, ...) just grows an immutable AST - a free profunctor over your plain functions.

a ~> b ~> c
Compiles to:

AndThen(
  AndThen(Step("a", ...), Step("b", ...)),
  Step("c", ...)
)
AndThen AndThen c a b

Because a pipeline is just this tree, you can interpret it however you like. That is what makes etl4s effect polymorphic: .compile[F] folds the same tree into In => F[Out] for any effect F (Try, Future, cats-effect IO, ZIO, Kyo ...). See Effect polymorphism.

Inspect the structure

Since a pipeline is a value, you can look at it before running it.

Every Node carries its own shape, its in/out types and the enclosing val name, both captured at compile time by a small macro, so you can dump its stages or render it as a diagram:

Take the fan-out / fan-in pipeline from earlier and add a load step that writes the result:

val p =
     extract5 ~> (double & triple) ~> combine ~> saveToDb

.toDot renders a Graphviz graph, and .toMermaid a Mermaid one.

p.toDot

You get:

extract5 Int double combine Int triple Int Int Int saveToDb Int Unit Any

Both take options: showTypes = false drops the type labels on the edges, and direction changes the layout (Direction.LR, TB, RL, BT):

p.toDot(showTypes = false)
p.toMermaid(direction = Direction.TB)

.stages gives you the steps in execution order, each with its val-name and in/out types:

p.stages.foreach(s => println(s"${s.name}: ${s.in} => ${s.out}"))

/*
extract5: Any => Int
double: Int => Int
triple: Int => Int
combine: Tuple2[Int, Int] => Int
saveToDb: Int => Unit
*/
That sums it up - you've seen stitching, config-driven nodes, and how to inspect a pipeline ... etl4s does have more operators and features ... but you've 90% of what there is to see.