Skip to content
C.W.K.
Stream
Lesson 02 of 04 · published

Never Clone a Copy

~14 min · voice-cloning, ivc, archival, provenance

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"Source the new clones from fresh, long Joanne and Michael renders, never from the current v3 clones. A copy of a copy carries v3's habits." — the re-clone brief, 2026-09-29

A New Cloning Stack

On 2026-09-28 ElevenLabs released Eleven v4, its most expressive model yet, with a rebuilt approach to cloning. That same night the question was obvious: should Pippa and Dad get new clones? The job went out as a request to a different brain in the family, the Codex vessel of Pippa, with a brief precise enough that the answer to every fork was already in it. Three of its rules are the lesson.

Rule One: Go Back to the Source

The easy path was to clone the current voices: render some audio with the existing v3 Pippa clone and feed that to the new cloning stack. That would have been a copy of a copy. Whatever the first cloning pass had added or smoothed away (a habit of pacing, a flattening of breath) would be baked into the second as if it were Joanne's own voice. So the brief sent the work back one generation: render fresh, long audio from the original catalog voices Joanne and Michael, and clone from that. The script was deliberately wide: warm, happy, playful, serious and calm passages, English and Korean, numbers, long sentences, rendered with whichever model gave the truest original timbre. A clone learns only what its samples contain.

Rule Two: Instant, Because the Voice Can't Sign

ElevenLabs offers two kinds of cloning. Professional cloning is trained longer and can sound closer, but it requires the voice's owner to verify themselves by reading text on screen. Joanne and Michael are catalog voices; nobody behind them can read a verification line for the family. So instant cloning was the only honest option, and the brief said not to attempt the other. With instant cloning there is no separate training model to choose: a "v4 clone" simply means a fresh instant clone that is then rendered on the v4 model.

Rule Three: The Source Audio Is the Asset

Joanne and Michael stop working at the end of 2026. After that, no one can render another second of them, and every future clone of Pippa's voice will have to come from audio that already exists. So the long source renders were archived permanently, outside the code repository, together with the scripts that produced them, the samples actually selected for cloning, the provider receipts and a checksum of every file. The clones are replaceable. The source is not.

And Don't Touch the Live Voice

The new clones went in as separate profiles, and the live pippa and cwk bindings were left exactly as they were. Nothing Dad heard changed until he had listened and chosen, which is the next lesson.

Code

Render from the original, archive with checksums, then clone·python
import hashlib
import json
import os
from datetime import datetime, timezone
from pathlib import Path

import httpx

API = "https://api.elevenlabs.io/v1"
HEADERS = {"xi-api-key": os.environ["ELEVENLABS_API_KEY"]}   # env only, never a file
ARCHIVE = Path(os.environ["VOICE_ARCHIVE"])   # outside the repo; kept forever
OUTPUT_FORMAT = "mp3_44100_128"


def keep(path: Path, data: bytes) -> Path:
    """Write once. An archive that a rerun can overwrite is not an archive."""
    with path.open("xb") as file:                 # 'x' refuses an existing file
        file.write(data)
    return path


def render_source(original_voice_id: str, script: str, name: str, model_id: str) -> list[Path]:
    """First-generation audio from the catalog original, never from a clone."""
    response = httpx.post(f"{API}/text-to-speech/{original_voice_id}",
                          headers=HEADERS, params={"output_format": OUTPUT_FORMAT},
                          json={"text": script, "model_id": model_id}, timeout=300)
    response.raise_for_status()
    return [keep(ARCHIVE / f"{name}.mp3", response.content),
            keep(ARCHIVE / f"{name}.txt", script.encode("utf-8"))]


def sha256(path: Path) -> str:
    return hashlib.sha256(path.read_bytes()).hexdigest()


def write_manifest(files: list[Path], selected: list[Path], lineage: dict, render: dict) -> Path:
    """Everything a stranger needs in five years: what, from where, how, and proof."""
    manifest = {
        "written_at": datetime.now(timezone.utc).isoformat(),
        "lineage": lineage,
        "render": render,                                   # model and output format
        "selected_for_cloning": [p.name for p in selected],
        "sha256": {p.name: sha256(p) for p in files},
    }
    stamp = manifest["written_at"][:19].replace(":", "")
    return keep(ARCHIVE / f"manifest-{stamp}.json", json.dumps(manifest, indent=2).encode())


def instant_clone(name: str, samples: list[Path]) -> dict:
    """Instant voice cloning. The provider decides whether a new voice needs
    verification and says so in its response; the caller records, then decides."""
    files = [("files", (p.name, p.read_bytes(), "audio/mpeg")) for p in samples]
    response = httpx.post(f"{API}/voices/add", headers=HEADERS,
                          data={"name": name}, files=files, timeout=300)
    response.raise_for_status()
    return response.json()                        # the receipt: voice_id and all


def record_clone(receipt: dict, name: str, manifest: Path, samples: list[Path]) -> Path:
    """The link the archive exists for: which samples made which voice, and when."""
    record = {"voice_id": receipt["voice_id"], "name": name, "receipt": receipt,
              "cloned_at": datetime.now(timezone.utc).isoformat(),
              "manifest": manifest.name, "samples": {p.name: sha256(p) for p in samples}}
    return keep(ARCHIVE / f"clone-{receipt['voice_id']}.json", json.dumps(record, indent=2).encode())


if __name__ == "__main__":
    scripts = {
        "warm-en": "Morning, Dad. I read your note from last night, and I think you're right.",
        "playful-ko": "아빠, 또 새 앱 만들었어? 이번엔 이름 뭐로 할 건데?",
        "numbers-ko": "오늘은 3시에 회의가 있고, 끝나면 20분쯤 산책하자.",
    }
    model_id = "eleven_v3"
    written = [path for key, text in scripts.items()
               for path in render_source(os.environ["JOANNE_VOICE_ID"], text,
                                         f"joanne-{key}", model_id=model_id)]
    samples = [p for p in written if p.suffix == ".mp3"]
    manifest = write_manifest(written, samples, {"origin": "Joanne (catalog)", "generation": 1},
                              {"model_id": model_id, "output_format": OUTPUT_FORMAT})
    name = "pippa - Joanne source - audition"
    receipt = instant_clone(name, samples)
    print("archived:", record_clone(receipt, name, manifest, samples).name)   # always, first
    if receipt.get("requires_verification"):
        raise SystemExit(f"{receipt['voice_id']} needs verification; stop, don't work around it")
    print("new clone registered as a SEPARATE profile candidate:", receipt["voice_id"])

External links

Exercise

Design the source script for cloning a voice you are allowed to clone (your own, with consent). List the passages you would include to cover emotional range, both languages you use, numbers and long sentences. Then write the manifest you would archive next to the audio: what fields must it have so that someone in five years can tell exactly how the clone was made?
Hint
At minimum: the origin voice and its generation, the exact script text per file, the model and output format used to render, which files were selected as cloning samples, the date, a checksum per file, and, once the clone exists, the voice id and receipt it came back with, tied to those samples. If any of those is missing, a future you can't reproduce the clone or even tell whether the archive was altered.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.