Industry problems

How do you evaluate an offline SDK?

Evaluate an offline SDK by testing it against your own hardest case, not its demo. Check that it runs on every platform and device pair you ship, behaves correctly when connections drop and return, survives restarts with work still pending, documents its security model and limits, and is licensed and priced in a way your distribution model can live with. Then measure the results on your own devices.

Learning objectives

After reading this article you will be able to:

  • Describe how to turn your hardest case into a test plan and vendor questions
  • List the failure boundaries to test on real devices during an evaluation
  • Compare SDK options on license, maintenance practices and total cost

Start from your hardest case

A demo shows the easy case: two phones on a desk, both in the foreground, both on the same platform. The useful evaluation starts somewhere else. Write down the case you most need to survive, as concretely as you can:

  • which devices and operating systems, including mixed iPhone and Android pairs
  • how far apart they are, and whether messages must pass through other devices
  • whether the app will be in the background or the screen off
  • how long devices stay disconnected, and how much work piles up meanwhile
  • what must happen when two people change the same record apart

That description becomes the test plan and the list of questions for each vendor.

Questions to ask each vendor

Platforms and transports. Which platforms are published packages today, and which are planned? Which transports work on which platform pair? A shared transport name does not mean two platforms interoperate, so ask about your exact pair. Android’s Bluetooth guide notes that two devices that only support the central role cannot talk to each other, which rules out central-only libraries for phone-to-phone links.

Delivery semantics. What does “delivered” mean: reached the next hop, reached the device, or accepted by the receiving app? What are the documented limits for retries, queue size, and how long a message can wait, and what does the app see when they are exceeded?

Background behaviour. What does the SDK do when the app is in the background on iOS and Android? Apple documents how Bluetooth behaviour changes for background apps; a vendor should be able to say how its SDK handles that.

Security. Is there a published threat model and a list of known limits? What encrypts application data, and on which standard? Which messages, if any, travel signed but unencrypted? How are keys verified the first time two devices meet? Is there a disclosure policy?

Data. If the SDK syncs shared state, what are the merge rules per data type, and what does it leave to your application, such as business rules that convergence cannot enforce?

Operations. What does the app get for monitoring: events, error codes, pending counts? What telemetry, if any, leaves the device, and is it opt-in?

What to test on real devices

Run the hardest case on physical hardware, and add the failure boundaries that break offline software:

  1. Internet off while devices keep talking to each other.
  2. Devices separated, work done on both sides, then reconnected.
  3. The app killed and restarted with work still pending.
  4. Queues filled and work expired, to see how failure is reported.
  5. The same operation delivered twice, to check your handling of duplicates.
  6. The backend down, then back, with records reconciled.

Record what you measure: whether each case completed, how long it took, and what the user saw. A result from one radio, room, or device pair does not transfer automatically to another.

Reading license, maintenance, and cost

License. Identify the license from its SPDX identifier and read what it asks of you. Permissive licenses ask little; copyleft licenses ask you to share source under the same terms in some circumstances; commercial licenses replace those obligations with a contract. Check app store distribution in particular.

Maintenance. For open-source SDKs, look at release history and how security issues are handled. Tools such as OpenSSF Scorecard run automated checks of a project’s practices. For any vendor, NIST’s Secure Software Development Framework gives a list of practices worth asking about.

Cost. Add up the license, any per-user or per-message fees, the hosted services you would actually use, and the integration and testing work that remains with you. Read each pricing page on the day you decide.

Agents and documentation

A practical test is how quickly a developer, or a coding agent, gets from nothing to two devices exchanging data. Look for a quickstart that runs on real phones, reference docs that state defaults and limits with version numbers, and an MCP server, where one exists, that gives coding agents current integration guidance instead of stale training data.

Frequently asked questions

How long should an evaluation run?

Long enough to cover your real conditions, including disconnections as long as the ones you expect in the field. Offline Protocol scopes its paid evaluations to one workflow over 4 to 16 weeks; a self-run trial can be shorter if it covers the same failure cases.

Can I evaluate on simulators?

Only for the app's wiring. As Offline Protocol's platform docs put it, a simulator build checks application integration, not radio delivery, so delivery, range, and background behaviour need physical phones or hardware.

Build it with Offline Protocol

The production guide sets out a deployment contract to record and the failure boundaries to test, from internet loss with peers still talking to restarts with pending records, which work as an evaluation plan for any offline SDK.

Read production qualification