etl4s
API Reference
  • .requires[T]
  • .provide(env)
  • .provideContext(env)
  • Reader[T, Node]
  • Etl4sCtx[T]

Configuration

When writing pipelines, you often need to:

.requires declares what a node needs. .provide supplies it once at the top. Config flows through automatically - it's just a Reader monad (Config => Node[In, Out]) with some syntax.

import etl4s._

case class Cfg(key: String)

val readData   = Node("data")
val tagWithKey = Node[String, String].requires[Cfg] { cfg => data =>
  s"${cfg.key}: $data"
}

val pipeline = 
     readData ~> tagWithKey

pipeline.provide(Cfg("secret")).unsafeRun()
You will get:
"secret: data"

Config stays out of your type signatures

Some effect systems bake config into the core data type as another type slot. For etl4s it would look like Node[Cfg, In, Out]

etl4s leaves the node at Node[In, Out]. .requires[Cfg] wraps just that node in the config it asks for (a Reader[Cfg, Node[In, Out]]).

The operators don't care which is which:

val load = Node[Unit, String](_ => "data")
val tag: Reader[Cfg, Node[String, String]] =
  Node[String, String].requires[Cfg] { c => s => c.key + s }

load ~> tag

~>, &, >>, etc compose plain and config-aware nodes together, infer the combined requirement, and leave you one .provide at the edge.

Config Propagation

Build modular configs with traits. etl4s infers what your pipeline needs:

trait HasDb { def dbUrl: String }
trait HasAuth { def apiKey: String }
case class AppConfig(dbUrl: String, apiKey: String) extends HasDb with HasAuth

val save = Node[String, Unit].requires[HasDb] { cfg => data =>
  println(s"Saving to ${cfg.dbUrl}: $data")
}

val fetch = Node[Unit, String].requires[HasAuth] { cfg => _ =>
  s"Fetched with ${cfg.apiKey}"
}

val toUpper = Node[String, String](_.toUpperCase)


val pipeline = 
     fetch ~> toUpper ~> save

pipeline.provide(AppConfig("jdbc:pg", "secret-key")).unsafeRun()

.requires[T] turns a node into a Reader[T, Node].

How config propagation works

The composition operators (~>, &, &>, >>) work directly on these config-aware nodes, and there are three rules that govern how requirements flow.

1. Plain nodes connect straight to config-aware ones. A plain Node requires nothing, so mixing it in adds no requirement. The pipeline still asks only for what the Reader nodes need:

val fetch   = Node[Unit, String].requires[HasAuth] { c => _ => s"got ${c.apiKey}" }
val toUpper = Node[String, String](_.toUpperCase)

val p = 
     fetch ~> toUpper

p.provide(AppConfig("jdbc:pg", "key")).unsafeRun()
HasAuth fetch ~> toUpper = HasAuth fetch ~> toUpper

2. Unrelated configs merge to an intersection. When two nodes require different configs, etl4s combines them to T1 & T2, here HasAuth & HasDb. You then .provide a single value that lives in the overlap (the AppConfig above, which extends both).

HasAuth HasDb

3. A subtype absorbs its supertype. If one node needs HasDb and another needs a subtype AppConfig <: HasDb, the merged requirement is just AppConfig. The more specific type (the smaller set) wins.

HasDb AppConfig

Etl4sCtx

Etl4sCtx[T] organizes config-driven nodes into modules:

case class DbConfig(url: String, timeout: Int)

object DataPipeline extends Etl4sCtx[DbConfig] {

  val fetch = Etl4sCtx.Extract[Unit, String] { cfg => _ =>
    s"Connected to ${cfg.url} with timeout ${cfg.timeout}s"
  }

  val save = Etl4sCtx.Load[String, Unit] { cfg => data =>
    println(s"Saving to ${cfg.url}: $data")
  }

  val pipeline = fetch ~> save
}

DataPipeline.pipeline.provide(DbConfig("jdbc:pg", 5000)).unsafeRun()

Scala 2

Use explicit types for better inference:

Transform.requires[Config, String, String] { cfg => input =>
  cfg.key + input
}
In Scala 3, the preferred syntax is:
Transform[String, String].requires[Config] { cfg => input =>
  cfg.key + input
}