Detecting structural gaps in a screenplay treatment is a different problem than summarizing text or generating dialogue. It requires a model to have a working theory of narrative mechanics, including what makes a story event structurally significant versus contextually relevant, how escalation pressure accumulates and dissipates across a three-act arc, and how the relationship between character interiority and external event creates the texture we recognize as dramatic structure.
We built the beat engine from scratch rather than fine-tuning a general-purpose language model for structural annotation, primarily because we found that general-purpose models with good summarization and generation capabilities still failed consistently at specific structural tasks. They could tell you what happened in a treatment. They could not reliably tell you whether what happened was doing the structural work it needed to do.
The Annotation Problem: What Are We Actually Labeling?
Before we could train a model to identify structural gaps, we had to agree on what a structural gap actually is. This turned out to be the hardest part of the first six months.
The naive definition is: a structural gap is a place where the standard beat framework says something should happen but the text does not show it happening. The problem with this definition is that it produces excessive false positives on non-standard structures, which includes a very high proportion of Indian film material. A model trained on this definition would flag as gaps many things that are deliberate structural choices.
The definition we ended up with is more functional: a structural gap is a place where the narrative requires a dramatic decision or escalation event to maintain forward momentum, but neither is present and the absence creates drag on the story's progression. This definition requires the model to have a theory of what creates narrative momentum, which is harder to specify but much more useful in practice.
Getting annotators to apply this definition consistently required extended calibration sessions. We would show two annotators the same treatment excerpt and ask them independently to identify structural gaps. Then we would compare their annotations and discuss the disagreements. The discussions were often more valuable than the annotations: they forced us to articulate intuitions that we had been applying inconsistently and to build explicit rules where we had only been pattern-matching.
Training on Indian Material: Coverage and Bias
We made the decision early on that our training data would be primarily Indian screenplay material. This was not obvious. There is more annotated screenplay data available in English from Hollywood sources, and bootstrapping from that corpus before adapting to Indian material would have been faster. We chose not to do it because we were concerned about structural biases that would be hard to remove once baked in.
Hollywood screenplay structure has particular signatures that are partly aesthetic and partly commercial: the tendency toward external-action-driven escalation in Act 2, the requirement for a clear protagonist want that drives each scene, the treatment of subplots as complications to the main plot rather than as co-equal narrative tracks. Indian commercial cinema has different signatures, and the structural logic of those differences matters for what counts as a gap versus what counts as a choice.
Our training corpus ended up being a mix of Hindi commercial cinema treatments, OTT series breakdowns from recent Indian productions, and a smaller set of regional language material. The Hindi commercial material has good coverage; the regional material is thinner, and our beat engine's reliability varies accordingly. We want to be honest about that. The engine is more confident on mainstream Hindi OTT material than on Tamil or Malayalam material, and users working on regional content should weight the feedback with that caveat in mind.
The Specificity Decision
When we designed the output format for beat annotations, we faced a choice between two approaches. The first was a fast, high-level output: a score for each act indicating overall structural health, plus flags at named beat positions (inciting incident, midpoint, etc.) where the standard placement was missing or misweighted. Quick to generate, easy to read, usable in five minutes.
The second approach was what we actually built: granular annotations at the scene or sequence level, each explaining not just where a gap exists but what narrative function is unfulfilled at that point and what kinds of story developments could address it. Much slower to generate, more expensive to compute, requires the user to actually read and think about the feedback.
We chose the second approach because our early testing showed that high-level structural scores had almost no effect on what writers did next. Writers who received a low midpoint score and nothing else did not know how to revise toward a better midpoint. Writers who received a note explaining that the midpoint sequence did not produce a meaningful reversal of the protagonist's external situation, which meant Act 2B started without the escalation pressure it needed to sustain forward momentum, had something to work with. The note told them what was wrong in terms they could act on.
The tradeoff is that the slower annotation cycle is genuinely slower. A full beat analysis on a feature treatment takes minutes rather than seconds. We have received feedback that this friction is real, particularly for writers who want quick iterative feedback during active drafting. We are working on a faster surface-level mode that can run on shorter text chunks during active writing, while reserving the full analysis for complete treatments and series breakdowns. The full analysis is not going away; the fast mode will complement it, not replace it.
What the Engine Cannot Do
The beat engine reads treatments and scene breakdowns well. It does not read formatted scripts well. Structural analysis on formatted script pages is a different technical problem: the signal-to-noise ratio is different, the structural information is distributed across action lines and dialogue rather than narrative prose, and the model needs to reconstruct the dramatic events from the screenplay representation rather than reading them directly.
We are working on formatted script analysis, but it is a separate engineering challenge and we are not close to shipping it. Writers who want beat analysis on their formatted drafts currently need to extract a prose summary of each scene and run the analysis against that. It is an extra step, and we know it is awkward.
The engine also does not evaluate dialogue quality, character consistency, or tonal coherence. These are adjacent problems that we want to address, and we have prototypes in internal testing. But they are not the same problem as structural analysis, and we did not want to ship them before we were confident they were adding value rather than noise.
The discipline we try to maintain is: do one thing well before adding the next thing. The beat engine does structural gap analysis on narrative treatments. It does that well enough that writers find it genuinely useful. That is the right foundation to build on.