Forum Discussion
Copilot Studio Agent Unable to Retrieve Usable DOCX/XLSX Content from OneDrive (and SharePoint)
hi edfencer The behavior described here points more toward binary content handling than an issue with the DOCX/XLSX files themselves.
The fact that test.txt comes back correctly, while the DOCX response starts with PK but contains “?” replacement characters, is a strong indication that the binary file is being interpreted as text somewhere in the Copilot Studio/Power Automate flow. Once the binary data has been decoded as text, the original bytes can be lost, which would explain why python-docx or openpyxl cannot open the reconstructed file.
A few things are worth checking:
- Confirm the output type of the OneDrive Get file content action. It should remain binary rather than being converted to a string.
- Avoid passing the file content through variables, JSON parsing, string manipulation, or other steps that may implicitly convert binary data to text.
- If the content needs to be passed to Python or another service, send it as binary or base64, rather than copying the displayed text representation.
- Test the same approach with a small DOCX/XLSX file to rule out payload-size or connector limitations.
- Check whether the same behavior occurs when using SharePoint - Get file content instead of OneDrive. If both behave the same way, that would further suggest an issue with how the flow/action is exposing binary content rather than with OneDrive itself.
One important point: seeing PK at the beginning doesn't necessarily mean the binary content is intact. The “?”characters are a good indication that some bytes have already been incorrectly decoded as Unicode, and those original byte values generally cannot be recovered reliably from that string.
If the requirement is for the agent to understand the contents of DOCX/XLSX files, another approach may be preferable to passing the raw Office file to Python. For example, extract the document text or Excel data first and pass the structured content to the agent. If the actual requirement is to preserve and process the original file, however, the binary/base64 representation needs to be maintained end-to-end.
It would also be useful to check the exact output schema returned by Get file content and how that output is being passed to the next action. That should help identify exactly where the binary-to-text conversion is occurring.
- edfencerAug 31, 2026Copper Contributor
I tried creating 4 test files in different format, i.e. txt, md, docx and pdf. Here's the results.
File Size Last modified Content retrieved test.txt 11 bytes 2026-08-26T11:17:31Z ✅ hello world test.md 11 bytes 2026-08-26T11:17:31Z ✅ hello world test.docx 13,973 bytes 2026-08-26T11:16:51Z ⚠️ raw binary only test.pdf 16,646 bytes 2026-08-31T04:18:16Z ⚠️ raw binary only I also tried SharePoint Get file content tool with these 4 files and received the same results. Both the OneDrive and SharePoint Get file content actions return binary file content. When these actions are exposed directly as Copilot Studio tools, the binary output is coerced into a Unicode text string rather than returned as a file resource or base64-encoded content. Invalid byte sequences are replaced with U+FFFD (�), irreversibly corrupting DOCX and PDF payloads.
- yonho00Aug 30, 2026Copper Contributor
Hello, I got stuck with the same problem described the above. Copilot Studio suggests me that I should ask Copilot Studio administrator to "enable a binary-preserving download path or expose infer_content_type as adjustable." Do you know if and how the above problem can be fixed on an Admin or agent maker level?