name: ppt-audio-to-video
description: Convert narration audio plus slide decks into a narrated video. Use when the user has an audio-only mp4/m4a/mp3/wav and a ppt/pptx/pdf deck, and needs slide images, transcript extraction, slide timing planning, or final mp4 rendering with whisper-cpp and ffmpeg.
Use this skill when the source video has narration audio but no usable slide visuals, and the final deliverable should be a slide-based lecture video.
Resolve bundled scripts relative to this skill directory. If the runtime has already opened this SKILL.md, prefer paths like scripts/extract_slide_outline.py and scripts/render_from_timing_csv.py instead of machine-specific absolute paths.
mp4/m4a/mp3/wav, ppt/pptx, pdf, and any pre-rendered slide images.Prefer an existing pdf or image directory for rendering. Treat pptx as the source of slide text and as a fallback for export.
Prepare tools.
7w4.net收录了海量优质技能插件。
ffmpeg, ffprobe, pdftoppm.whisper-cli from whisper-cpp plus a multilingual model such as ggml-small.bin.If only pptx exists and no pdf/images exist, prefer Keynote or PowerPoint export on macOS. Use soffice only as fallback because profile or rendering issues are common.
Produce slide images.
pdf exists, render it to images:
bash
pdftoppm -png -r 200 "$PDF" "$OUTDIR/slide"pptx exists, export to pdf or slide images with Keynote or PowerPoint, then continue from pdf.Keep slide filenames ordered and stable, such as slide-01.png, slide-02.png, ...
Extract slide text.
bash
python3 scripts/extract_slide_outline.py \
--pptx "$PPTX" \
--out "$WORKDIR/slide_outline.csv"Use the output to identify slide titles, distinctive keywords, and section changes.
Extract clean audio for ASR.
mp4, extract mono wav:
bash
ffmpeg -y -i "$AUDIO_MP4" -ar 16000 -ac 1 -c:a pcm_s16le "$WORKDIR/audio.wav"If the source is already wav/mp3/m4a, convert to the same mono wav form if needed.
Transcribe with whisper-cli.
bash
whisper-cli -ng \
-m "$MODEL" \
-f "$WORKDIR/audio.wav" \
-l zh \
-ocsv -osrt -of "$WORKDIR/transcript"transcript.csv for downstream parsing. transcript.srt is useful for manual review.If GPU allocation fails on macOS, retry with -ng to force CPU mode.
Build slide_timings.csv.
csv
slide,start_sec,end_sec,duration_sec,reason
1,0.000,15.000,15.000,opening title and agenda
2,15.000,100.000,85.000,architecture overview starts hereduration_sec = end_sec - start_sec.Validate that the last end_sec matches the audio duration or is within a small tolerance.
Render the final video.
bash
python3 scripts/render_from_timing_csv.py \
--images "$SLIDE_IMAGES_DIR" \
--timings "$WORKDIR/slide_timings.csv" \
--audio "$WORKDIR/audio.wav" \
--output "$OUT_VIDEO"The script generates an ffconcat file, validates timing continuity, and calls ffmpeg to encode the final mp4.
Verify and iterate.
ffprobe.slide_timings.csv and rerun the render script.slide_timings.csv.Install dependencies on macOS if missing:
brew install ffmpeg poppler whisper-cpp
Typical multilingual model download:
mkdir -p .models
curl -L 'https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-small.bin' -o .models/ggml-small.bin
scripts/extract_slide_outline.py
Extract slide text from pptx into CSV or JSON for timing analysis.scripts/render_from_timing_csv.py
Validate a timing CSV, generate an ffconcat, and render the final video with ffmpeg.这个工具整体质量不错,能较好地将音频旁白与幻灯片合成视频。文档说明详细,流程清晰,两个脚本运行稳定。主要优点是工作流完整、验证机制可靠。不足之处是使用前需要手动安装多个外部工具和 Python 库,对新手不太友好;另外有些步骤需要手动执行,自动化程度可以进一步提高。