Mugafi started as a question: could a software system give a screenwriter the same quality of structural feedback that a good script editor gives, at the speed that actual writing requires? Not a summarization tool or a grammar checker. A system that reads narrative structure the way an experienced reader reads narrative structure, identifying where the story is doing the work it needs to do and where it is not.
We started building in early 2023, and the honest account of that first year is a mix of things that worked, things we got wrong and had to redo, and one decision that we nearly made but did not make, which turned out to be the most important decision of the whole project.
What We Tried First and Why It Did Not Work
The tempting approach was to fine-tune a large language model on a curated dataset of screenplays and generate structural feedback by prompting the fine-tuned model with a treatment. We tried a version of this. The output was fluent and confident and often plausible-sounding. It was also unreliable in ways that were difficult to predict or fix.
The fundamental problem was that fine-tuning on screenplay text produces a model that is good at generating screenplay-like text, not good at analyzing structural function. Structural analysis requires a different kind of reasoning: it requires the system to hold a theory of narrative mechanics and evaluate the text against that theory, rather than predicting what text typically comes next in screenplay documents.
We also found that the screenplay corpus we had access to was heavily biased toward Hollywood material. A model fine-tuned on that corpus would consistently rate treatments as structurally deficient if they did not conform to the three-act Hollywood template, which is not a useful evaluation for Indian material.
By mid-2023 we had decided to build the structural analysis layer as a separate system rather than as a prompted language model, with explicit representations of structural concepts rather than implicit representations learned from text.
The Decision That Changed the Architecture
The important decision I mentioned was whether to build a system that returns prescriptive structural feedback (your inciting incident should be here, your midpoint is placed wrong, here is what should happen) or a system that returns diagnostic structural feedback (the narrative logic here requires a decision or escalation event, and the text does not provide one).
We almost built the prescriptive version because it felt more tractable: you could define the expected beat positions for a given format, check whether the treatment had events at those positions, and generate notes telling the writer what to add or move. It would have been faster to build and faster to generate output.
What stopped us was a series of conversations with working Indian screenwriters who were very clear about something: they did not want to be told where their story should turn. They wanted to understand why a structural problem existed in terms of narrative logic, not in terms of template compliance. A note that says "your midpoint is at page sixty-five instead of page fifty-five" is useless unless the writer understands what narrative function the midpoint is supposed to serve and why that function is not being served at page sixty-five. The position is a symptom; the function is the diagnosis.
We rebuilt the annotation system around functional descriptions of structural elements. What does this beat need to accomplish? What narrative conditions need to be in place for the story to move forward from this point? When those conditions are not present in the text, what kind of events could satisfy them?
Building the Annotation Framework
Translating that conceptual decision into an annotation framework that humans could apply consistently was about four months of iterative work. We worked with a small group of script editors and senior screenwriters who had extensive experience with Indian commercial cinema and OTT material. The sessions were structured as annotated disagreements: we would bring ambiguous examples and ask each annotator to explain their structural intuition in as much detail as they could, then use the disagreements as the basis for refining the annotation guidelines.
The hardest category was what we started calling "functional equivalence": cases where a treatment achieved the same narrative function through a structurally unconventional route. A midpoint reversal does not have to be an external event; it can be an internal realization that changes the protagonist's understanding of their situation completely, with no visible external change. How do you reliably distinguish between a midpoint that is absent and a midpoint that is present but handled in an unconventional way?
Our resolution was to annotate for functional outcome rather than event type. The annotation asks: does the protagonist's situation or understanding change significantly at this narrative position, in a way that changes the problem they are trying to solve in Act 2B? If yes, the midpoint function is satisfied regardless of how the change occurs. If no, it is flagged as a gap.
What Surprised Us in Training
Two things surprised us significantly during the training process.
The first was how much the structural patterns in Indian commercial cinema depend on the interval. The interval, the explicit narrative break at roughly the midpoint of a Hindi film, is not just a theatrical convention; it has structural function. The scenes immediately before and after the interval are doing specific narrative work that is different from what happens at the midpoint in a Western three-act structure. Our initial annotation framework treated it as a simple position marker and produced systematically wrong gap detections for those sequences. We had to build interval-awareness into the structural logic explicitly.
The second surprise was how unreliable our initial coverage of ensemble-driven structures was. Many Indian commercial films have multiple protagonists with co-equal narrative weight. Our single-protagonist structural model applied to those films produced annotations that were confusing at best. Structural gap detection for ensemble narratives requires tracking multiple arc progressions simultaneously, and the gap analysis has to evaluate whether the narrative obligations of the ensemble as a unit are being met, not just whether any single character's arc is progressing.
Both of these required rebuilding sections of the system rather than just adjusting parameters. The interval work in particular took several weeks. But getting them right was the difference between a system that would generate plausible-sounding but systematically wrong output for a significant portion of Indian material and a system that could handle those structures appropriately.
Where We Are Now
The beat engine we shipped is not the system we originally planned. It is more careful about what it claims to know, more explicit about its coverage limitations, and more granular in its output than what we initially designed. It is also slower than we wanted it to be, which is a tradeoff we made deliberately and would make again.
The thing we got right from the beginning: building the system specifically for Indian material rather than adapting a general-purpose tool. Every architectural decision we made that was rooted in the specific characteristics of Indian film structure paid off. Every place where we borrowed from a Western screenplay analysis framework uncritically had to be revised.
We are still building. The character consistency tracking, the dialogue analysis, the series-level structural view across multiple episodes are all in development. Each of those is a different problem from beat detection, and each will require the same kind of ground-up work on Indian material that the beat engine required. But we know the approach that works, and we know what the mistakes look like before we make them. That is the benefit of getting the first system right the hard way.