The engineering moved faster than the commercial thinking.
A critique of 61 graduation projects from the 2025/26 professional training programmes, the catalogue they sit in, and the freelancing outcomes reported alongside them. Every figure here can be opened down to the slide it came from.
Six branches did not submit
Assiut, Banha, Mansoura, New Capital, New Valley and Port Said are absent from every figure in this report. They represent 23% of ITP offerings and 15% of PTP. Read every headline as 16 of 22 branches reporting.
Everything on this page opens. Click any dot, bar, cell or row for the record behind it · Ctrl+K jumps to any project by name · P fills the screen with one exhibit at a time.
The finding
What the review concluded, and how it got there.
How we judged
Two scores, kept deliberately apart. Quality asks how good the thinking is. Documentation asks how much of it was written down. Collapsing them would have ranked branches on slide design, so nothing in this report does.
What each criterion feeds
Click any criterion for the scoring anchors as they were applied. The four AI criteria are scored on every project and shown in every record — they are simply not averaged into the quality score.
The exhibit above is the whole method in one picture, and the part worth pausing on is the middle column. 4 criteria are scored but held out of the quality score: AI depth, agentic depth built, agentic depth needed, and cost against revenue model. A project that builds an impressive agent should not out-rank one that thought harder about the problem it was solving, and averaging those four in is precisely what would have made it. They drive their own findings instead — the built-against-needed comparison, and the cost analysis — and they appear in full in every project record.
Not stated is never a low score
This is the rule the whole two-axis design rests on, so it is stated once, here. Not statedNE means the source material did not say. A 1 means stated, and bad. Conflating them would punish branches for thin documentation rather than thin thinking, so a not-stated criterion is left out of the quality average and counted on the documentation axis instead. Inference is not evidence: if a target market can only be reverse-engineered from a solution sentence, it is not stated. Across the portfolio 461 criteria came back not stated, and an automated audit found zero cases of one being scored as a 1 instead.
What we found
The headline is not the one we expected going in. We assumed this review would mostly document AI-washing — projects claiming agentic AI that turned out to be chatbots. It did not.
Twenty-one of 61 projects claim agentic AI. Twenty of them can back the claim with a described architecture: named agents, tool use, plan-act-observe loops. Students across the network can genuinely build this.
The weakness is on the other side of the ledger. 14 of 61 projects state how they would make money. 2 state how they would reach a customer. 35 are entering markets our research rated crowded or saturated, and only 9 address a gap incumbents have actually left open.
The nine measures
Every tile opens the exhibit behind it. Figures cover 16 of 22 branches.
The AI is real — and it is all the same AI
We scored two things separately and deliberately: what was actually built, and what the problem genuinely required — judged from the problem alone, before looking at the build. Comparing them gives four groups.
What was built against what was needed
Every dot is one project; click it to open the full record. Cells lay their projects out as a small grid, so counts stay exact where projects share coordinates. Use the legend to isolate a group.
Only 2 projects built more AI than their problem needed. That is a good result, and it contradicts the common assumption that students bolt agents onto everything. The 7 that needed more than they built are the ones worth acting on — real problems where the team reached for a chatbot when the problem justified more.
The seven that needed more than they built
| Project | Branch | Needed | Built |
|---|---|---|---|
| DAR Platform | Damanhour | 3/5 | 1/4 |
| Hiring Stage | Damanhour | 3/5 | 1/4 |
| NetFix AI | Ismailia | 3/5 | 1/4 |
| NutriScan AI | Smart Village | 3/5 | 1/4 |
| Training Simulation | Smart Village | 3/5 | 1/4 |
| Mongez | Smart Village | 3/5 | 1/4 |
| MedAgents | Sohag | 3/5 | 1/4 |
These are teachable, not failures. Each already has a problem worth solving, with the validation done — which makes them ready-made briefs for the next intake.
How the four groups are derived
C3 agentic depth is scored from the described architecture only — the word “agent” in a bullet is not evidence, a plan→act→observe loop with tool calls is. C4 agentic necessity is scored from the problem statement alone, before looking at what was built. The group is then mechanical: built-more-than-needed is C3 ≥ 2 with C4 ≤ 2; needed-more-than-built is C3 ≤ 1 with C4 ≥ 3.
One template, applied regardless of the problem
The concerning reading is not that RAG is popular. It is that RAG appears on problems that do not obviously need retrieval. RAG appears in 24 projects across 13 branches and LLM in 14 across 8. Behind them sits the same supporting cast — Qdrant, Ollama, Semantic Kernel, Gemini — in near-identical combinations. Two Damanhour projects list identical AI stacks.
What the portfolio is built from
Counted across the 45 projects that named a stack at all. Click any bar to list the projects behind it. How these counts were derived.
The business thinking is not
The same teams that can build a multi-agent system cannot say who pays, how much, or how the product reaches anyone — and almost none of them checked who was already there. This is the largest single finding in the review.
Every project, by what it says about money
One cell per project. 14 of 61 say how they would make money; 2 also say how they would reach a customer; 6 give enough for Whether the per-customer numbers work to be assessed at all. Click any cell for the record.
Two projects in 61 described a route to a customer. That is not a student failing — it is not currently being asked of them. A one-page business canvas as a mandatory graduation artefact would close most of this gap, and it is a week of teaching against a weakness that affects 77% of projects.
Nobody checked who else was building this
35 of 61 projects entered markets our research rated crowded or saturated. Only 9 addressed a gap incumbents had genuinely left open.
How contested the markets entered are
Scored per project against named global leaders and Egyptian players, citing 426 sources across 180 comparables.
Five projects landed on top of a funded incumbent
| Project | Branch | Incumbent | Why it matters |
|---|---|---|---|
| Furnora | Damietta | Homzmart | $15M Series A, already operating in Damietta furniture |
| Elshamy Pharmacies | Smart Village | Chefaa | 1.5M+ monthly users; El Ezaby runs 185+ branches |
| Brixel | Minia | Kuadra | Egyptian AI construction platform, launched 13 July 2026 — days before submission |
| On-Demand Home Nurse | Smart Village | 7keema | billed as Egypt's first on-demand home nursing app |
| CLINIKA | Ismailia | AI4Docs.AI | Arabic ambient clinical scribe built in Cairo |
Two hours of searching at proposal stage would have redirected several of these toward genuine whitespace. Open any project to read the full comparable set, including pricing and where the incumbent falls short.
The expensive ones cannot pay for themselves
Every project was modelled at a common 1,000 monthly usersMAU so the figures stay comparable. These are order-of-magnitude estimates with stated assumptions, not budgets.
Every project by estimated monthly AI cost
Bars are coloured by whether the project says how it would make money. The pattern at the top of the chart is the finding: the most expensive projects are disproportionately the ones with no stated way to pay for themselves. Open any bar for its full assumption set. The median above excludes the 15 projects that describe no AI at all and therefore cost nothing to run — why.
The five above $1,000 per month
All five are driven by speech-to-text or multimodal processing on clinical and media workloads. Two carry an explicit caveat that 1,000 monthly users is artificial for a B2B clinician tool — surfaced in the record rather than hidden.
Documentation varies more than quality
The two axes from the opening section, drawn against each other. This is the exhibit that decides which projects can be ranked at all — and it is the reason the ranking that follows covers 8 projects rather than 61.
Quality against documentation, every project
One dot per project, coloured by band. The trend line is fitted to these dots rather than drawn from the quoted figure, so the two cannot disagree. Click any dot for the record.
Documentation bands
| Band | Projects | What it means |
|---|---|---|
Fully documentedband A | 8 | Scored on 12 or more of the 15 quality criteria. The only projects rankable with confidence. |
Partly documentedband B | 30 | Enough to assess, not enough to rank against a fully documented project. |
Barely documentedband C | 23 | Too little written down to apply most of the rubric. Not a judgement on the project. |
23 of 61 projects are barely documented. Compare within a band, not across one.
A caveat we are not dropping
Quality score correlates with documentation score at +0.78. Two readings are available and both are probably partly true: teams that thought harder produced better projects and better documentation; and our score partly rewards documentation depth regardless of quality. That is why projects are banded, and why the ranking below covers only the fully documented ones.
Where the thin documentation is
Average documentation score by branch
Smart Village submitted the most projects (19) and documented them the least, averaging 1.86 against 3.40 across every other branch. 16 of its 19 could not be assessed on enough criteria to rank fairly.
This is not evidence that Smart Village's projects are weak. It is evidence that we cannot tell — which, for the branch running 51% of PTP Intake 46, is its own problem. Its game and data-management projects arrived as feature tag-clouds with no problem statement, customer or business model. Alexandria's entire deck yielded 36 characters of machine-readable text. Damanhour and Qena submitted structured canvases. All were responding to the same request.
Ranked — the projects we can rank
8 projects had enough documentation to score on 12 or more of the 15 quality criteria. These are the only projects that can be ranked with confidence.
Fully documented projects, ranked by quality
| # | Project | Branch | Quality | Criteria | Documentation | AI group |
|---|---|---|---|---|---|---|
| 1 | Arena — AI-Powered Gym Management System | Qena | 3.86 | 14 | 4.33 | |
| 4 | ShopBrain Analytics | Sohag | 3.58 | 12 | 3.67 | |
| 9 | DAR Platform | Damanhour | 3.43 | 14 | 4.33 | |
| 11 | Smart Feedback | Aswan | 3.36 | 14 | 4.00 | |
| 16 | Hiring Stage | Damanhour | 3.23 | 13 | 4.33 | |
| 21 | TaskPilot | Aswan | 3.08 | 12 | 4.00 | |
| 33 | AURA AI | Ismailia | 2.83 | 12 | 3.67 | |
| 38 | Brixel | Minia | 2.79 | 14 | 3.67 |
Quality is the mean of the 15 quality criteria that could be applied, leaving out anything not stated. “Criteria” is how many of the 15 were applied — the number that makes the quality score trustworthy or not.
Arena — AI-Powered Gym Management System leads on a full 14 criteria — an AI gym management system whose deterministic clinical-safety layer validates AI-generated workout and nutrition plans against ACSM/AHA/ADA guidelines, a guard none of its three comparables implements.
The other 53 projects are not ranked here
They are all in the explorer in Act 2, with every score, justification and comparable intact. They are unranked because their documentation could not support a rank, which is a different statement from scoring badly.
What to do about it
Five changes, each answering a finding above. None of them needs a new course — four are a week of teaching or a change to what is asked for at submission.
Make a one-page business canvas a mandatory graduation artefact
Who pays, how much, how they are reached, and what it costs to serve them. One page, submitted with the project.
Because 14 of 61 projects say how they would make money and 2 say how they would reach a customer. It is not currently being asked of them.
Check who else is building it before approving a project, not after
Two hours of searching at proposal stage, recorded and reviewed with the supervisor — not a post-hoc justification after the build.
Because 35 of 61 projects entered markets we rated crowded or saturated, and five landed directly on top of a funded incumbent.
Teach the choice, not just the tool
Ask every team to justify why their AI approach fits their problem, and to name the simpler approach they rejected. Include at least one project brief where the right answer is deterministic code.
Because RAG and the same four supporting tools appear across the portfolio on problems that do not obviously need retrieval. Two projects list identical stacks.
Teach cost modelling alongside AI architecture
A team that can estimate its own inference bill will design differently, and will find the constraint before a customer does.
Because The most expensive projects to run are disproportionately the ones with no stated way to pay for themselves.
Reissue the under-built projects as briefs for the next intake
Each already has a real problem, a named affected group, and the validation done. What they lack is the architecture the problem deserved.
Because 7 projects have a problem that justified an agent and a build that stopped at a chatbot. The hard part of the brief is already written.
Set a documentation floor before the next review
Half the findings above are weaker than they should be because the submissions vary so much. A short required template — problem, customer, revenue, stack, team, demo link — would cost a branch an afternoon and would make the next intake's projects comparable to each other for the first time.
Explore
The same data, unsummarised. Filter it yourself.
Every project
All 61 projects, with every score, justification, comparable and cost assumption intact. Filter, sort, then open any row for the complete record.
The full portfolio
Sort by any column. Quality is only comparable within a documentation band — the “of 15” column tells you how much of the rubric could actually be applied.
Unattributed records are kept, not dropped
1 project carries no branch and 1 carries no programme; 38 of 61 carry no track, because most decks never printed one. These appear under an “unattributed” option in each filter rather than disappearing from the counts. Track recovery went from 4 to 23 during the image pass; the rest genuinely are not in the source.
By branch
16 branches submitted. Each row opens that branch's full record — its projects, its rollups, and its freelancing rows.
Reporting branches
Averages cover only that branch's own projects. Branches submitting one or two projects have averages that should not be read as a ranking.
The six that did not submit
These branches are absent from every figure in this report. Between them they run the offerings listed here, which is the size of the blind spot.
Why PTP cannot be compared across branches
PTP Intake 46 runs in only 11 of 22 branches, and 28 of its 55 tracks — 51% — sit at Smart Village. Any cross-branch PTP comparison is close to meaningless; use department and area instead.
Reference
The catalogue, the raw freelancing rows, and the method.
Freelancing — reported, not yet reconciled
The freelancing figures branches reported alongside their projects. This is raw data with its quality problems left visible: the reconciliation against the system, and the job-type and platform analysis, are a separate piece of work that has not been done.
Read this before quoting any number here
These 119 rows are what branches stated. Every row is marked derived: N — nothing here was computed by us. 12 rows fail an internal consistency check and 3 income figures are statistical outliers, all listed below rather than quietly cleaned. Rows also differ in What each row countsgrain — some count a track, some a branch, some a person — so they must not be summed across grains or compared against a per-track system figure.
Known data problems
Recorded during the merge and carried through unchanged. Resolving these is the first step of the reconciliation, not something to do silently afterwards — reconciling against known-bad numbers wastes the exercise.
Freelancing summary — every reported row
What each row counts is shown on every row. Where a track name could not be matched to the catalogue, the raw string as written on the slide is shown instead.
Named freelancers
38 individuals with a recorded project. Platform is missing for 11 of them. Open a row for the project idea, tools and client.
What still has to happen
Obtain the system export and confirm what its rows count and which columns it carries; resolve the arithmetic and percentage failures above; join on offering_id and produce a variance table of reported against system; and only then analyse job types and platforms. Zagazig's single “Total $35,857” must be split by programme first.
The programme catalogue
The cleaned reference the whole report joins against: 22 units and 219 tracks at branchesoffering, keyed by offering_id (programme × department × track × branch).
Offerings
Both programmes in scope: PTP Intake 46, and ITP 25/26 with rounds 1 and 2 merged.
Source files
The 22 branch and department files this report was built from. Much of the content was locked inside slide images and was recovered by re-reading every referenced slide at 150 DPI.
Open gaps
What would have to be requested to close the record. Blocking gaps are the six missing branch reports.
Method, verification and limits
How the scores were produced, what was done to check them, and what this report cannot tell you. None of this is buried — the caveats are part of the finding. The scoring itself is explained in How we judged.
The rubric
27 criterion identifiers, of which 23 carry a score and 4 are structured fields rather than scores — C2 records the agentic claim as written, C5 derives the group, C6 flags AI-washing, and C7 holds the cost model. Of the 23 scored, 15 set the quality score, 3 set the documentation score, 4 are scored but held out of both, and E2 is scoreable in principle but was never filled — the cross-portfolio clustering it depends on was not run, so it is not stated on all 61 projects. Not every criterion applies to every project: the number actually applied ranges from 21 to 23, and a record distinguishes scored, not stated and not applied as three different states.
Every criterion, with its anchors
Expand any criterion for the full scoring anchors as they were applied.
How the technology counts were derived
Whole-word matching, on purpose
A term counts once per project where it appears as a whole word in the technologies or AI-components text as submitted. Substring matching on a word boundary is deliberate: the source cells concatenate tools, so Qdrant(RAG) and chatbot – Qdrant – RAG are both genuine RAG mentions that a delimiter split would miss. The companion curriculum document reports lower counts because it split on delimiters only — these figures are the more complete of the two.
What the median AI cost is taken over
Excluding the projects that run no AI
15 of the 61 projects describe no AI at all (C1 = 0) and therefore cost nothing to run. They are excluded from the median, which is taken across the 46 projects that do use AI — including them would drag it to $200 and describe a “median AI running cost” using projects that run no AI. The portfolio total keeps all 61.
Verification
An independent agent re-scored a stratified sample of nine projects spanning the full documentation range, without sight of the original scores. Across 180 comparable score pairs: 73% exact agreement, 90% within one point, and 8 of 9 groups agreed. Group agreement matters most, since the built-against-needed comparison is this report's central claim — and the verifier independently reproduced TAGER as having built more than it needed, the counterintuitive call in the original pass.
The automated not-stated audit found zero violations: no criterion was scored 1 where the underlying source field was empty, which is the exact failure the rubric was written to prevent.
An inconsistency we found and did not hide
The blind re-score disagreed on whether a criterion was not stated or scoreable in 16 of 180 pairs (9%), one-directionally: on thin projects the original pass scored E3 and D3 numerically where the verifier judged them not stated. Quality was recomputed with the disputed criteria removed. The mean shift was +0.00, the largest single move was +0.38, and no project changed band. The scores were therefore retained and the finding recorded here rather than silently corrected.
Limits
What this report cannot tell you
| Limit | What it means |
|---|---|
| Six branches did not submit | Assiut, Banha, Mansoura, New Capital, New Valley, Port Said — 23% of ITP offerings and 15% of PTP. |
| Quality correlates with documentation at +0.78 | Projects are banded and must be compared within a band. Stated openly rather than dropped. |
| 38 of 61 projects have no track | Most decks never printed one. Department-level analysis is therefore partial. |
| Prior art is research, not audit | 180 comparables across 426 sources, gathered per project. Search quality is recorded per project and shown in each record. |
| AI costs are order-of-magnitude | All modelled at a common 1,000 monthly users so they stay comparable. Two carry an explicit caveat that 1,000 is artificial for a B2B clinician tool. |
Every one of these is live. None is a reason to discard the findings; all are reasons to read them at the right resolution.