Documents submitted from SharePoint drive automated routing and storage in the data lake. Tagging quality directly affects whether the right extraction pipeline runs and whether downstream reporting is trustworthy.
For reports, Doc Type is the primary routing key. For invoices (Doc Type = Invoice), routing uses Vendor plus Invoice Category when required.
Please ensure every file is correctly categorized before submission:
- Project Category
- Places the work in the correct program bucket for reporting and prioritization.
- Vendor
- Must match the controlled list — free text is not supported. Incorrect vendor values can cause rejection or mis-routing.
- Doc Type
- This is the extraction routing key: it tells the platform which specialized process to run. Choosing the wrong doc type means the wrong automation may run on your file.
- Invoice Category
-
Expected on invoice files (Doc Type = Invoice). When a vendor has more
than one automated invoice extraction, Invoice Category tells the
platform which process to run — select the value from SharePoint's
controlled list.
- Fiserv invoices — Invoice Category required
- CSI invoices — Invoice Category required
- Extract No Database
-
Use this when you have a document that does not match a
known automated Doc Type, or when you need a quick generic read of a PDF
or image without running a production validated extraction flow.
- When to use: Unknown layouts, one-off files, or
documents waiting for a dedicated extractor. Pick the choice that
best describes the document:
- Invoice Extraction — PDF or image invoices that do not match a mapped Vendor + Doc Type = Invoice route (unknown vendor or layout).
- Report Extraction or Other Extraction — non-invoice PDFs or images that need a quick generic read without a production validated report flow.
- How to use: Set Extract No Database on the SharePoint item, complete Project Category, and submit from this portal. Vendor and Doc Type are optional when this field is populated (you may still fill them in for filing if helpful).
- What you get — Invoice Extraction: Same as Report/Other for now — an Excel workbook Extraction_Generic_Data_Extraction_No_DB_Write_<timestamp>.xlsx (generic OCR). The dedicated unknown-invoice CSV/validation pipeline is not live in staging yet. Not written to the PRI database. When a dedicated invoice extractor exists for your vendor, use Vendor + Doc Type = Invoice (and Invoice Category when required) instead.
- What you get — Report Extraction / Other Extraction: An Excel workbook named Extraction_Generic_Data_Extraction_No_DB_Write_<timestamp>.xlsx returned to the library. Each source file becomes a worksheet with extracted tables and text blocks. This is a generic OCR output only — not written to the PRI database and not a substitute for a completed, validated extraction flow when one exists for your Doc Type.
- Supported files: PDF, JPEG, PNG, and TIFF (all three choices).
- When to use: Unknown layouts, one-off files, or
documents waiting for a dedicated extractor. Pick the choice that
best describes the document:
- True Up
- Use this field consistently so files can be filtered and reviewed correctly. It is used for filtering even when it is not required for a file to appear in the picker.