tinygrad/examples/compile_efficientnet.py

from pathlib import Path
from extra.models.efficientnet import EfficientNet
from tinygrad.tensor import Tensor
from tinygrad.nn.state import safe_save
from extra.export_model import export_model
from tinygrad.helpers import getenv, fetch
import ast

if __name__ == "__main__":
  model = EfficientNet(0)
  model.load_from_pretrained()
  mode = "clang" if getenv("CLANG", "") != "" else "webgpu" if getenv("WEBGPU", "") != "" else "webgl" if getenv("WEBGL", "") != "" else ""
  prg, inp_sizes, out_sizes, state = export_model(model, mode, Tensor.randn(1,3,224,224))
  dirname = Path(__file__).parent
  if getenv("CLANG", "") == "":
    safe_save(state, (dirname / "net.safetensors").as_posix())
    ext = "js" if getenv("WEBGPU", "") != "" or getenv("WEBGL", "") != "" else "json"
    with open(dirname / f"net.{ext}", "w") as text_file:
      text_file.write(prg)
  else:
    cprog = [prg]
    # image library!
    cprog += ["#define STB_IMAGE_IMPLEMENTATION", fetch("https://raw.githubusercontent.com/nothings/stb/master/stb_image.h").read_text().replace("half", "_half")]

    # imagenet labels, move to datasets?
    lbls = ast.literal_eval(fetch("https://gist.githubusercontent.com/yrevar/942d3a0ac09ec9e5eb3a/raw/238f720ff059c1f82f368259d1ca4ffa5dd8f9f5/imagenet1000_clsidx_to_labels.txt").read_text())
    lbls = ['"'+lbls[i]+'"' for i in range(1000)]
    inputs = "\n".join([f"float {inp}[{inp_size}];" for inp,inp_size in inp_sizes.items()])
    outputs = "\n".join([f"float {out}[{out_size}];" for out,out_size in out_sizes.items()])
    cprog.append(f"char *lbls[] = {{{','.join(lbls)}}};")
    cprog.append(inputs)
    cprog.append(outputs)

    # buffers (empty + weights)
    cprog.append("""
  int main(int argc, char* argv[]) {
    int DEBUG = getenv("DEBUG") != NULL ? atoi(getenv("DEBUG")) : 0;
    int X=0, Y=0, chan=0;
    stbi_uc *image = (argc > 1) ? stbi_load(argv[1], &X, &Y, &chan, 3) : stbi_load_from_file(stdin, &X, &Y, &chan, 3);
    assert(image != NULL);
    if (DEBUG) printf("loaded image %dx%d channels %d\\n", X, Y, chan);
    assert(chan == 3);
    // resize to input[1,3,224,224] and rescale
    for (int y = 0; y < 224; y++) {
      for (int x = 0; x < 224; x++) {
        // get sample position
        int tx = (x/224.)*X;
        int ty = (y/224.)*Y;
        for (int c = 0; c < 3; c++) {
          input0[c*224*224 + y*224 + x] = (image[ty*X*chan + tx*chan + c] / 255.0 - 0.45) / 0.225;
        }
      }
    }
    net(input0, output0);
    float best = -INFINITY;
    int best_idx = -1;
    for (int i = 0; i < 1000; i++) {
      if (output0[i] > best) {
        best = output0[i];
        best_idx = i;
      }
    }
    if (DEBUG) printf("category : %d (%s) with %f\\n", best_idx, lbls[best_idx], best);
    else printf("%s\\n", lbls[best_idx]);
  }""")

    # CLANG=1 python3 examples/compile_efficientnet.py | clang -O2 -lm -x c - -o recognize && DEBUG=1 time ./recognize docs/showcase/stable_diffusion_by_tinygrad.jpg
    # category : 281 (tabby, tabby cat) with 9.452788
    print('\n'.join(cprog))
[ready] Replacing os with pathlib (#1708) * replace os.path with pathlib * safe convert dirnames to pathlib * replace all os.path.join * fix cuda error * change main chunk * Reviewer fixes * fix vgg * Fixed everything * Final fixes * ensure consistency * Change all parent.parent... to parents 2023-08-31 01:41:08 +08:00			`from pathlib import Path`
move things, clean up extra (#2292) * move things * idk why pylint needs that now * delete unused 2023-11-14 12:18:40 +08:00			`from extra.models.efficientnet import EfficientNet`
clang backend (#572) * start clang backend * mostly working * no group for reduce w clang * it compiles * compiles * a11y * minor fixups * formatting * add a test * rename test 2023-02-21 10:18:18 +08:00			`from tinygrad.tensor import Tensor`
move state to nn/state (#1619) 2023-08-22 22:36:24 +08:00			`from tinygrad.nn.state import safe_save`
simple exporting models (#1344) * unified exporting * json exporting * ignore more * simplified buffer export * added dtypes * added assert * swift example * fix tests * linter * remove whitespace * fixed tests * remove swift example * remove unintended changes * allow callable models to be used * whitespace * more readable json export * name change * whitespace * whitespace 2023-08-02 00:35:48 +08:00			`from extra.export_model import export_model`
add name support to fetch (#2407) * add name support * use fetch in gpt2 * remove requests from main lib, networkx also optional * umm, keep that assert * updates to fetch * i love the walrus so much * stop bundling mnist with tinygrad * err, https * download cache names * add DOWNLOAD_CACHE_VERSION * need env. * ugh, wrong path * replace get_child 2023-11-24 06:16:17 +08:00			`from tinygrad.helpers import getenv, fetch`
[ready] Replacing os with pathlib (#1708) * replace os.path with pathlib * safe convert dirnames to pathlib * replace all os.path.join * fix cuda error * change main chunk * Reviewer fixes * fix vgg * Fixed everything * Final fixes * ensure consistency * Change all parent.parent... to parents 2023-08-31 01:41:08 +08:00			`import ast`
compile_tensorflow: add initialize and tests 2023-02-23 12:50:53 +08:00
Webgpu support (#1077) * initial commit * 81 passing * 105 passing tests * 148 passing * CI tests * install dep on ci * try opencl pkgs * try using vulkan * down to only 6 failing * refactor * cleaning up * another test skipped due to buffer limit * linter * segfault * indent fix * another segfault found * small touchups * Fix max and maxpool tests * Add constant folding * Add javascript export script * better asserts in codegen * manual upcasting * reverted token type change * skip safetensor test due to unsupported type * FIx efficientnet and all other model tests * Remove np copy * fixed indent and missing import * manually destroy the buffer * revert back to length * linter errors * removed extra val * skip broken tests * skipping more tests * Make the page pretty * Save model weights as safetensor * Fix imagenet to c test * Fix second imagenet to c bug * Async and paralel kernel compilation * workgroup support * reversed local size * fixed non local bug * correct local groups * ci experiment * removed typo * Fix define local by using shared memory * Refactor * try running on mac * match metal tests * add more workers * scope down tests * trying windows runner * fixed windows env * see how many it can do * merged master * refactor * missed refactor * increase test suite coverage * missing import * whitespace in test_efficientnet.py * getting there * fixed reset * fixed bufs * switched to cstyle * cleanup * min/max rename * one more linter issue * fixed demo * linter * testing ci chrome * add unsafe webgpu arg * add build step * remove WEBGPU from cmd line * use module * try forcing directx * trying forced metal backend * temp disable conv2d for CI * disable conv_trasnpose2d --------- Co-authored-by: 0x4d - Martin Loretz <20306567+martinloretzzz@users.noreply.github.com> Co-authored-by: George Hotz <72895+geohot@users.noreply.github.com> 2023-07-13 03:52:06 +08:00			`if __name__ == "__main__":`
			`model = EfficientNet(0)`
			`model.load_from_pretrained()`
webgl backend in extra (#3041) * WebGL WIP * 84% of ops passing test * tests passing 100% * Cleanup, refactor * Shave off some lines * Work on dtypes * TestOps at 100% again * Efficient net shaders compile in browser webgl2 * Compile all efficientnet shaders in browser * Create empty textures for tensor buffers * Run program. Up next weight loading * Exported WebGL model working * Add tests, refactor * Explicit cast alu for GLSL * Fix CI tests * WebGL efficientnet demo * Compile and run yolov8 in browser * Fix imports * Simplify yolo compile * Fix boolbool and cast cmplt to float More tests * Do std tests pass on CI? * Skip std tests on CI * Remove explicit_cast_alu hack, and solve it in code_for_op * Move to new dtype-less alloc api * Remove local size hack: optimize local_size only if device has local * Remove glsl.py, and move content to cstyle * dont_use_locals in opts * Fix dtype tests * type_map in CStyleLanguage * Make core changes smaller, cleaner, refactor export_model and demo * Skip pad_slice * Simplify: render_const, render_conditional * solve bool alu for other binops, cleaner ops_webgl * Fix noopt hack * Remove some skipIfs * WebGL image hack * type_names is a better name * global_max * Fix dtype import * Fix type_names -> type_map * Fix lint * Remove webgpu, back to 5k lines (#3040) * remove webgpu * max 5000 lines * revert those to master * retain that cstyle --------- Co-authored-by: Ahmed Harmouche <ahmedharmouche92@gmail.com> 2024-01-09 01:29:13 +08:00			`mode = "clang" if getenv("CLANG", "") != "" else "webgpu" if getenv("WEBGPU", "") != "" else "webgl" if getenv("WEBGL", "") != "" else ""`
Enable Multi-Output Export (#2179) * Enable Multi-Output Export * Add test * Update examples and lint * fix padding * test ops * dummy commit to rerun test * revert cuda lint * Enforce tuple/list of tensors * subscripted generics * put back webgpu test * Re-enable WebGPU Efficientnet test 2023-10-31 09:42:26 +08:00			`prg, inp_sizes, out_sizes, state = export_model(model, mode, Tensor.randn(1,3,224,224))`
[ready] Replacing os with pathlib (#1708) * replace os.path with pathlib * safe convert dirnames to pathlib * replace all os.path.join * fix cuda error * change main chunk * Reviewer fixes * fix vgg * Fixed everything * Final fixes * ensure consistency * Change all parent.parent... to parents 2023-08-31 01:41:08 +08:00			`dirname = Path(__file__).parent`
simple exporting models (#1344) * unified exporting * json exporting * ignore more * simplified buffer export * added dtypes * added assert * swift example * fix tests * linter * remove whitespace * fixed tests * remove swift example * remove unintended changes * allow callable models to be used * whitespace * more readable json export * name change * whitespace * whitespace 2023-08-02 00:35:48 +08:00			`if getenv("CLANG", "") == "":`
[ready] Replacing os with pathlib (#1708) * replace os.path with pathlib * safe convert dirnames to pathlib * replace all os.path.join * fix cuda error * change main chunk * Reviewer fixes * fix vgg * Fixed everything * Final fixes * ensure consistency * Change all parent.parent... to parents 2023-08-31 01:41:08 +08:00			`safe_save(state, (dirname / "net.safetensors").as_posix())`
webgl backend in extra (#3041) * WebGL WIP * 84% of ops passing test * tests passing 100% * Cleanup, refactor * Shave off some lines * Work on dtypes * TestOps at 100% again * Efficient net shaders compile in browser webgl2 * Compile all efficientnet shaders in browser * Create empty textures for tensor buffers * Run program. Up next weight loading * Exported WebGL model working * Add tests, refactor * Explicit cast alu for GLSL * Fix CI tests * WebGL efficientnet demo * Compile and run yolov8 in browser * Fix imports * Simplify yolo compile * Fix boolbool and cast cmplt to float More tests * Do std tests pass on CI? * Skip std tests on CI * Remove explicit_cast_alu hack, and solve it in code_for_op * Move to new dtype-less alloc api * Remove local size hack: optimize local_size only if device has local * Remove glsl.py, and move content to cstyle * dont_use_locals in opts * Fix dtype tests * type_map in CStyleLanguage * Make core changes smaller, cleaner, refactor export_model and demo * Skip pad_slice * Simplify: render_const, render_conditional * solve bool alu for other binops, cleaner ops_webgl * Fix noopt hack * Remove some skipIfs * WebGL image hack * type_names is a better name * global_max * Fix dtype import * Fix type_names -> type_map * Fix lint * Remove webgpu, back to 5k lines (#3040) * remove webgpu * max 5000 lines * revert those to master * retain that cstyle --------- Co-authored-by: Ahmed Harmouche <ahmedharmouche92@gmail.com> 2024-01-09 01:29:13 +08:00			`ext = "js" if getenv("WEBGPU", "") != "" or getenv("WEBGL", "") != "" else "json"`
[ready] Replacing os with pathlib (#1708) * replace os.path with pathlib * safe convert dirnames to pathlib * replace all os.path.join * fix cuda error * change main chunk * Reviewer fixes * fix vgg * Fixed everything * Final fixes * ensure consistency * Change all parent.parent... to parents 2023-08-31 01:41:08 +08:00			`with open(dirname / f"net.{ext}", "w") as text_file:`
simple exporting models (#1344) * unified exporting * json exporting * ignore more * simplified buffer export * added dtypes * added assert * swift example * fix tests * linter * remove whitespace * fixed tests * remove swift example * remove unintended changes * allow callable models to be used * whitespace * more readable json export * name change * whitespace * whitespace 2023-08-02 00:35:48 +08:00			`text_file.write(prg)`
			`else:`
			`cprog = [prg]`
			`# image library!`
add name support to fetch (#2407) * add name support * use fetch in gpt2 * remove requests from main lib, networkx also optional * umm, keep that assert * updates to fetch * i love the walrus so much * stop bundling mnist with tinygrad * err, https * download cache names * add DOWNLOAD_CACHE_VERSION * need env. * ugh, wrong path * replace get_child 2023-11-24 06:16:17 +08:00			`cprog += ["#define STB_IMAGE_IMPLEMENTATION", fetch("https://raw.githubusercontent.com/nothings/stb/master/stb_image.h").read_text().replace("half", "_half")]`
simple exporting models (#1344) * unified exporting * json exporting * ignore more * simplified buffer export * added dtypes * added assert * swift example * fix tests * linter * remove whitespace * fixed tests * remove swift example * remove unintended changes * allow callable models to be used * whitespace * more readable json export * name change * whitespace * whitespace 2023-08-02 00:35:48 +08:00
			`# imagenet labels, move to datasets?`
add name support to fetch (#2407) * add name support * use fetch in gpt2 * remove requests from main lib, networkx also optional * umm, keep that assert * updates to fetch * i love the walrus so much * stop bundling mnist with tinygrad * err, https * download cache names * add DOWNLOAD_CACHE_VERSION * need env. * ugh, wrong path * replace get_child 2023-11-24 06:16:17 +08:00			`lbls = ast.literal_eval(fetch("https://gist.githubusercontent.com/yrevar/942d3a0ac09ec9e5eb3a/raw/238f720ff059c1f82f368259d1ca4ffa5dd8f9f5/imagenet1000_clsidx_to_labels.txt").read_text())`
simple exporting models (#1344) * unified exporting * json exporting * ignore more * simplified buffer export * added dtypes * added assert * swift example * fix tests * linter * remove whitespace * fixed tests * remove swift example * remove unintended changes * allow callable models to be used * whitespace * more readable json export * name change * whitespace * whitespace 2023-08-02 00:35:48 +08:00			`lbls = ['"'+lbls[i]+'"' for i in range(1000)]`
Allow multi-input model export (#1995) * Allow multi-input model export * Add model export unit test * Fix efficientnet compilation * Only run model export test on JIT supported devices * Skip export model test if not EXPORT_SUPPORTED_DEVICE 2023-10-07 19:13:34 +08:00			`inputs = "\n".join([f"float {inp}[{inp_size}];" for inp,inp_size in inp_sizes.items()])`
Enable Multi-Output Export (#2179) * Enable Multi-Output Export * Add test * Update examples and lint * fix padding * test ops * dummy commit to rerun test * revert cuda lint * Enforce tuple/list of tensors * subscripted generics * put back webgpu test * Re-enable WebGPU Efficientnet test 2023-10-31 09:42:26 +08:00			`outputs = "\n".join([f"float {out}[{out_size}];" for out,out_size in out_sizes.items()])`
simple exporting models (#1344) * unified exporting * json exporting * ignore more * simplified buffer export * added dtypes * added assert * swift example * fix tests * linter * remove whitespace * fixed tests * remove swift example * remove unintended changes * allow callable models to be used * whitespace * more readable json export * name change * whitespace * whitespace 2023-08-02 00:35:48 +08:00			`cprog.append(f"char *lbls[] = {{{','.join(lbls)}}};")`
Allow multi-input model export (#1995) * Allow multi-input model export * Add model export unit test * Fix efficientnet compilation * Only run model export test on JIT supported devices * Skip export model test if not EXPORT_SUPPORTED_DEVICE 2023-10-07 19:13:34 +08:00			`cprog.append(inputs)`
Enable Multi-Output Export (#2179) * Enable Multi-Output Export * Add test * Update examples and lint * fix padding * test ops * dummy commit to rerun test * revert cuda lint * Enforce tuple/list of tensors * subscripted generics * put back webgpu test * Re-enable WebGPU Efficientnet test 2023-10-31 09:42:26 +08:00			`cprog.append(outputs)`
simple exporting models (#1344) * unified exporting * json exporting * ignore more * simplified buffer export * added dtypes * added assert * swift example * fix tests * linter * remove whitespace * fixed tests * remove swift example * remove unintended changes * allow callable models to be used * whitespace * more readable json export * name change * whitespace * whitespace 2023-08-02 00:35:48 +08:00
			`# buffers (empty + weights)`
			`cprog.append("""`
			`int main(int argc, char* argv[]) {`
			`int DEBUG = getenv("DEBUG") != NULL ? atoi(getenv("DEBUG")) : 0;`
			`int X=0, Y=0, chan=0;`
			`stbi_uc *image = (argc > 1) ? stbi_load(argv[1], &X, &Y, &chan, 3) : stbi_load_from_file(stdin, &X, &Y, &chan, 3);`
			`assert(image != NULL);`
			`if (DEBUG) printf("loaded image %dx%d channels %d\\n", X, Y, chan);`
			`assert(chan == 3);`
			`// resize to input[1,3,224,224] and rescale`
			`for (int y = 0; y < 224; y++) {`
			`for (int x = 0; x < 224; x++) {`
			`// get sample position`
			`int tx = (x/224.)*X;`
			`int ty = (y/224.)*Y;`
			`for (int c = 0; c < 3; c++) {`
Allow multi-input model export (#1995) * Allow multi-input model export * Add model export unit test * Fix efficientnet compilation * Only run model export test on JIT supported devices * Skip export model test if not EXPORT_SUPPORTED_DEVICE 2023-10-07 19:13:34 +08:00			`input0[c224224 + y224 + x] = (image[tyXchan + txchan + c] / 255.0 - 0.45) / 0.225;`
simple exporting models (#1344) * unified exporting * json exporting * ignore more * simplified buffer export * added dtypes * added assert * swift example * fix tests * linter * remove whitespace * fixed tests * remove swift example * remove unintended changes * allow callable models to be used * whitespace * more readable json export * name change * whitespace * whitespace 2023-08-02 00:35:48 +08:00			`}`
clang backend (#572) * start clang backend * mostly working * no group for reduce w clang * it compiles * compiles * a11y * minor fixups * formatting * add a test * rename test 2023-02-21 10:18:18 +08:00			`}`
			`}`
Enable Multi-Output Export (#2179) * Enable Multi-Output Export * Add test * Update examples and lint * fix padding * test ops * dummy commit to rerun test * revert cuda lint * Enforce tuple/list of tensors * subscripted generics * put back webgpu test * Re-enable WebGPU Efficientnet test 2023-10-31 09:42:26 +08:00			`net(input0, output0);`
simple exporting models (#1344) * unified exporting * json exporting * ignore more * simplified buffer export * added dtypes * added assert * swift example * fix tests * linter * remove whitespace * fixed tests * remove swift example * remove unintended changes * allow callable models to be used * whitespace * more readable json export * name change * whitespace * whitespace 2023-08-02 00:35:48 +08:00			`float best = -INFINITY;`
			`int best_idx = -1;`
			`for (int i = 0; i < 1000; i++) {`
Enable Multi-Output Export (#2179) * Enable Multi-Output Export * Add test * Update examples and lint * fix padding * test ops * dummy commit to rerun test * revert cuda lint * Enforce tuple/list of tensors * subscripted generics * put back webgpu test * Re-enable WebGPU Efficientnet test 2023-10-31 09:42:26 +08:00			`if (output0[i] > best) {`
			`best = output0[i];`
simple exporting models (#1344) * unified exporting * json exporting * ignore more * simplified buffer export * added dtypes * added assert * swift example * fix tests * linter * remove whitespace * fixed tests * remove swift example * remove unintended changes * allow callable models to be used * whitespace * more readable json export * name change * whitespace * whitespace 2023-08-02 00:35:48 +08:00			`best_idx = i;`
			`}`
clang backend (#572) * start clang backend * mostly working * no group for reduce w clang * it compiles * compiles * a11y * minor fixups * formatting * add a test * rename test 2023-02-21 10:18:18 +08:00			`}`
simple exporting models (#1344) * unified exporting * json exporting * ignore more * simplified buffer export * added dtypes * added assert * swift example * fix tests * linter * remove whitespace * fixed tests * remove swift example * remove unintended changes * allow callable models to be used * whitespace * more readable json export * name change * whitespace * whitespace 2023-08-02 00:35:48 +08:00			`if (DEBUG) printf("category : %d (%s) with %f\\n", best_idx, lbls[best_idx], best);`
			`else printf("%s\\n", lbls[best_idx]);`
			`}""")`
clang backend (#572) * start clang backend * mostly working * no group for reduce w clang * it compiles * compiles * a11y * minor fixups * formatting * add a test * rename test 2023-02-21 10:18:18 +08:00
simple exporting models (#1344) * unified exporting * json exporting * ignore more * simplified buffer export * added dtypes * added assert * swift example * fix tests * linter * remove whitespace * fixed tests * remove swift example * remove unintended changes * allow callable models to be used * whitespace * more readable json export * name change * whitespace * whitespace 2023-08-02 00:35:48 +08:00			`# CLANG=1 python3 examples/compile_efficientnet.py \| clang -O2 -lm -x c - -o recognize && DEBUG=1 time ./recognize docs/showcase/stable_diffusion_by_tinygrad.jpg`
			`# category : 281 (tabby, tabby cat) with 9.452788`
			`print('\n'.join(cprog))`