PencilHB builds SteelBox, an AI steel takeoff tool, so read this with that in mind. The criteria below are the ones we think any tool, including ours, should be held to.
The promise and the risk
A manual steel takeoff is mostly searching: finding members, finding the right sheet, reading labels, measuring, and typing. That part is repetitive and AI is good at it. The rest is judgement: is that label a W16×26 or a W16×36 under the leader line, does this bay really match the one next to it, is this member structural or miscellaneous. That part carries the risk, and AI is not good enough at it to be left alone.
So the question when evaluating AI takeoff software is not "how smart is the model" but "how does the tool handle the gap between what the model suggests and what I am willing to sign". The eight criteria below are about that gap.
1. Every suggestion shows its evidence
When the tool says a member is a W16×26, you should be able to see the region of the drawing it read, the confidence it has, and why it flagged the item for review if it did. A number without a source cannot be checked, and a takeoff that cannot be checked is not a takeoff. Be wary of tools that present a finished spreadsheet with no way back to the drawing.
2. Nothing becomes project data without a person confirming it
Look for an explicit approval gate. AI proposes, a person accepts, edits or rejects, and only then does the item count. This should apply to members, section labels, lengths and groups. Tools that auto-commit and let you "fix it later" push the checking work to the end, where it is slowest and most likely to be skipped under bid pressure.
3. Measurement is calibrated, never assumed
PDF drawings are routinely reissued at the wrong scale. A trustworthy tool asks you to confirm one known dimension on the sheet before it measures anything, and makes the calibration visible. If a tool gives you lengths without ever asking about scale, ask where they came from.
4. Grouping is safe
Applying one measured length to all members of the same section is a huge time saver and a classic source of error, because same section does not mean same length. The tool should suggest groups, let you approve which members join, and never overwrite a value you entered by hand. Conflicts should be shown, not silently resolved.
5. The drawing stays in front of you
You will spend the review looking at the drawing, not at a table. The workspace should keep the sheet large, jump to the next unresolved member automatically, and keep queue and detail panels out of the way. If the demo is mostly a spreadsheet, the tool was designed around the output, not the review.
6. It handles your drawing standards
Section notation differs between AISC (W16×26), Indian Standard (ISMB 300), Eurocode (IPE 300) and others, and so do sheet conventions and units. Ask which standards the tool has actually been tested on, in metric and imperial, and try it on your own sheets rather than the vendor's sample set.
7. It exports something you can use
The output needs to fit your estimating and detailing workflow: a member list with mark, section, count, length, weight and source, in a format your pricing tool accepts. Ask whether the export carries the review status of each item so a reviewer can see what was confirmed, edited or still open.
8. It is clear about your data
Drawings are client-confidential. Ask where they are stored, who can see them, whether they are used to train models, and how they are deleted. A vendor that cannot answer these questions in writing is not ready for your projects.
Questions to ask a vendor
- Show me the source region for this suggestion.
- What happens to an AI suggestion nobody has confirmed? Does it appear in the export?
- How is scale set, and can I see the calibration later?
- If I measure one beam and apply it to a group, what stops it overwriting a member I measured by hand?
- Which drawing standards and units have you tested on? Can I run it on my own set today?
- What does the export contain, and does it include review status and source?
- Where are my drawings stored, and are they used for training?
- What accuracy do you claim, and on what drawings was it measured?
Running a fair pilot
- Use real drawings you have already taken off, so you have a checked answer to compare against. Two or three sets of different quality is enough.
- Measure the things that matter: members found versus missed, wrong sections, wrong lengths, time from upload to a reviewed list, and how many items could be traced to a source.
- Include the reviewer's time. A tool that finds everything but takes as long to check as a manual takeoff has not saved anything.
- Try to break it. Give it a poorly scanned sheet, a crowded detail, and a plan with labels crossed by leader lines. How it fails matters more than how it succeeds.
- Keep the engineer in the loop. The person who signs the estimate should be the one confirming members, at least for the pilot.
FAQ
Can AI do a steel takeoff automatically?
It can find members, read labels and suggest lengths. Drawings are ambiguous and a takeoff carries commercial and safety risk, so reliable tools treat AI output as suggestions a person confirms, and show the evidence behind each one.
How accurate is AI steel takeoff software?
It depends on the drawings and the tool, and no vendor can honestly promise one number that holds across all projects. Test it on your own drawings and compare against a takeoff you have already checked.
What should I test in a pilot?
Members found versus missed, wrong labels, time to a reviewed list, and whether every number can be traced back to the drawing.
How SteelBox approaches this. SteelBox is built around the approval gate: it finds members, labels and marks, shows the source region and confidence for each, guides calibration and measurement, suggests groups you approve, and never turns a suggestion into project data on its own. It is in a limited early-access pilot for structural steel teams. Request early access, or start with how to do a steel takeoff from drawings.